NeFut Logo NeFut
中 Admin Login

[CS.AI] MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

In recent years, machine learning models have been deployed in high‑stakes domains, yet their inner workings remain opaque to practitioners. Conventional post‑hoc explanation methods demand expert skills: navigating high‑dimensional outputs, selecting the most appropriate explanations, and synthesizing evidence across disparate tools. To eliminate this knowledge barrier, we introduce MEA, a multi‑agent framework.

MEA consists of two agents. The Proposer automatically selects and configures explanation tools based on the question type and data modality (tabular, text, vision). The Actor is trained end‑to‑end with a faithfulness reward, converting tool outputs into natural‑language explanations that are grounded in model behavior.

We define three question categories—feature attribution, counterfactual reasoning, and spurious feature detection—each paired with a perturbation‑based faithfulness metric. Experiments reveal that state‑of‑the‑art LLMs often produce unfaithful explanations when left unconstrained. By augmenting the reward with a modality‑adaptive penalty, MEA consistently outperforms traditional post‑hoc explainers, agentic systems, and closed‑source baselines across six datasets, achieving faithfulness gains of +28% (tabular), +21% (text), and +34% (vision) over an untrained backbone.

More broadly, our results suggest that AI agents themselves can serve as scalable, adaptable interfaces for ML explainability, paving the way for natural‑language explanations that go beyond the fixed, single‑purpose tools that have long dominated the field.

Review

Original Source: https://arxiv.org/abs/2610.02480

[h] Back to Home