NeFut Logo NeFut
Admin Login

[CS.AI] From Plausible to Actionable: A Position on LLM Self‑Explanations

Published at: 2026-09-11 22:00 Last updated: 2026-09-12 06:35
#Machine Learning #LLM #Artificial Intelligence

Large language models (LLMs) can produce natural‑language explanations that justify their own decisions, a capability known as self‑explanation.

Self‑explanations have emerged as a promising avenue for explainable AI (XAI), especially for probing LLM behavior.

Yet these explanations often look plausible without guaranteeing that they faithfully mirror the model’s internal reasoning, an open question.

This opinion argues that self‑explanations can be highly plausible, questionably faithful, and nevertheless highly actionable.

From a conventional XAI standpoint, we point out limitations of current evaluation protocols for LLM‑generated self‑explanations, such as reliance on subjective human judgments and lack of checks on reasoning pathways.

We propose practical guidelines: (1) assess plausibility by measuring consistency between model outputs and explanations; (2) test faithfulness using adversarial inputs or causal interventions.

Moreover, evaluation should extend beyond plausibility and faithfulness to actionability—whether the explanation enables stakeholders to make informed decisions and take appropriate actions.

Example applications include medical diagnosis assistance, legal document review, and business decision support, illustrating how LLM rationalization can drive concrete actions.

In sum, we urge researchers to design evaluation frameworks that jointly consider plausibility, faithfulness, and actionability to advance the deployment of self‑explanations in real‑world systems.

Review

Original Source: https://arxiv.org/abs/2607.15957

[h] Back to Home