ReLiveGym is a diagnostic evaluation suite for long‑lived tasks, replaying weeks of real‑world news, market and social‑media streams in chronological order. Agents act sparsely on these streams, and the tasks vary in time‑sensitivity, reasoning intensity and recurrence. We benchmark eight base language models, examining how model choice and harness design—particularly the mechanism that decides when to act—affect performance. Results show that action‑timing is a crucial design axis for long‑lived tasks, and the optimal design differs across tasks and sometimes across models. We also enable continuous learning from hindsight feedback, which improves performance and mitigates observed failure modes. These findings suggest that model selection, timing mechanisms, and feedback utilization are key considerations when building unattended LLM agents. The implementation is available at https://github.com/SaharaLabsAI/ReLiveGym.
Review: This study offers a comprehensive benchmark for persistent agents, highlighting timing control and feedback‑driven adaptation as essential factors for real‑world deployment.