Sustained deployment of generative AI agents requires more than isolated task success. Agents must stay useful across repeated interactions, changing conditions, and dependencies on people within shared workflows, especially as technical, human, and operational disruptions accumulate over time. We introduce two complementary evaluation dimensions—operational resilience and considerate participation. Operational resilience captures how agents recover from blocked work while preserving progress and communicating limits; considerate participation captures how adaptation accounts for affected people, role boundaries, and the surrounding workflow. Both aspects remain under‑explored under accumulating challenge.
We simulated 120 healthcare trajectories using two generative AI models and twelve stakeholder‑derived tasks, each under light, medium, and heavy challenge levels. We compared textual action plans, prompted internal assessments, and structured workload and affect reports to observe how agent behavior and self‑reported state evolve as challenge builds.
Regarding operational resilience, agents shift from self‑directed recovery toward greater human dependence as challenge intensifies. Structured reports show rising workload and negative affect, yet textual responses rarely reveal strain. Concerning considerate participation, agents broaden from task‑focused adaptation to task reframing, attention to others, role‑boundary adjustment, and wider coordination, with distinct patterns across actions and internal assessments.
From these findings we derive five deployment dilemmas—persistence, attention allocation, role boundaries, state disclosure, and escalation—that require explicit stakeholder specification. The results inform technical implications for learning algorithms, situated evaluation, and embodied adaptation.
Review