MeshHeal is a fully decentralized self‑healing framework that addresses gray failures in decentralized LLM‑based multi‑agent systems. A gray failure occurs when an agent stays responsive but its problem‑solving quality continuously degrades, lacking enough evidence to change routing while still needing to rejoin after recovery.
At the fast timescale, MeshHeal builds an adaptive hierarchy. A single reviewer scores an output; if the score is uncertain or below a threshold, the case is escalated to a committee deliberation, and the output can be corrected before use.
At the slow timescale, a task‑ and ability‑conditioned peer‑relative detector aggregates scores across agents to separate persistent degradation from normal output variance. Persistent degradation triggers mandatory committee review and eventually excludes the degraded agent from ordinary routing; periodic recovery probes supply fresh evidence to reintegrate recovered agents.
To faithfully evaluate routing, we introduce Model‑Backed MAS Evaluation, which ties ability assignments to execution models, preventing routing errors that hidden when using only prompt‑based assignments.
Experiments on BBH, MATH, and MMLU‑Pro show that MeshHeal attains 0.839 accuracy in the degraded phase with only 51 k model tokens per task, compared to the strongest baseline Symphony’s 0.807 accuracy using 115 k tokens. Under staggered degradation and recovery, MeshHeal isolates degraded agents, keeps them out of normal task execution until they recover, and then restores them to standard routing.
Review: MeshHeal’s two‑timescale review and detection pipeline provides timely detection and correction of gray failures, achieving higher robustness with lower token overhead for decentralized LLM agent networks.