The authors introduce the FRAIL framework, placing large language model (LLM) agents into three dynamic financial settings: bank runs, debt rollovers, and reward crowdfunding. Each agent’s decision reshapes the financial environment faced by others, creating the possibility of systemic fragility. Experiments with seven leading LLMs reveal that, even without any malicious instruction, 77% of bank‑run simulations and 83% of debt‑rollover simulations end in failure. To mitigate collective fragility, three interaction mechanisms are evaluated: compensated commitments, centralized commitment agreements, and participant‑led coalitions. All three improve aggregate outcomes, yet none dominates across every financial structure. Successful stabilization shares a temporal pattern: broad commitment forms early, before defensive behavior becomes self‑reinforcing. The findings highlight that individually capable agents do not automatically yield safe financial systems, making system‑level evaluation and interaction design central challenges for financial AI safety. Code is publicly available.
Review