Large language models (LLMs) are increasingly deployed as multi‑turn agents that must sustain goals, use tools, and adapt to other agents over extended interactions. Existing work lacks auditable, multi‑turn, multi‑factor experiments that quantify LLM behavior under explicit constraints, as well as time‑resolved statistics that reveal how behavior unfolds over long horizons. To fill this gap we introduce a micro‑benchmark inspired by the Stanford marshmallow experiment. ReAct agents operate minute‑by‑minute, equipped with a “raise a question” tool and a per‑step budget, while we factorially manipulate social context (broadcast vs. isolation), personas (age, hedonic drive), and metacognitive policy (mandatory vs. optional tool use).
We evaluate 19,200 agent trajectories across 64 cells using Kaplan‑Meier survival curves and discrete‑time hazard models. Results show a sharp early “eat now” impulse, with only 75.9% of agents persisting to the end. The hazard model indicates isolation reduces per‑minute risk relative to broadcast, whereas a must‑use self‑questioning policy increases risk. On average agents ask $\approx 7.12$ questions and hit the per‑step budget in about $6\%$ of minutes. Questioning declines faster under broadcast than isolation.
Ablation experiments reveal that removing hedonic drive and/or age increases survival and narrows the broadcast/isolated gap, while the must‑vs‑may ordering remains unchanged. The combined ablation (no hedonic + no age) yields the highest completion rate, approaching $1.0$. These findings establish delay‑of‑gratification as a compact multi‑turn interaction benchmark that captures social contagion and tool‑use dynamics, providing a reproducible testbed and statistics for analyzing long‑horizon, multi‑agent behavior.
Review