Long-horizon LLM agents must preserve task state across extended sequences of observations, actions, tool calls, and intermediate beliefs. We examine these trajectories through three dynamical lenses: self-organized criticality, weak chaos, and metastable belief dynamics. Our framework aligns agent-implied states with benchmark-grounded states and quantifies stress accumulation, error avalanches, temporal dependence, local‑global mismatch, bounded divergence, belief‑basin transitions, and finite-size scaling under explicit null models.
Across 22 experiments—ranging from controlled puzzles and tool use to embodied tasks, multi‑hop retrieval, general‑assistant reasoning, and the Game of Life—we find that locally valid actions can persist after global state fidelity breaks down; stress can trigger abrupt collapse; error sequences exhibit long memory; dependency depth alters the propagation regime; and larger horizons accommodate larger avalanches. At the same time, divergence remains bounded, belief states are metastable rather than fully chaotic, and there is no support for universal power laws, critical points, or shared intervention optima.
These results suggest that a science of agent world models should rely on trajectory‑level dynamical diagnostics instead of terminal reward alone.
Review