NeFut Logo NeFut
Admin Login

[CS.AI] Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?

Published at: 2026-09-08 22:00 Last updated: 2026-09-09 09:08
#AI #Machine Learning #LLM

AI oversight typically relies on ground‑truth answers to validate models, yet what counts as appropriate AI behavior is itself contested. Consequently, evaluations of moral reasoning in large language models (LLMs) often sidestep realistic ambiguity. To address this, we introduce a standard that remains operative under ambiguity: the structural quality of a model’s defence for its verdicts. The defence is measured via a four‑phase dialectical protocol grounded in Walton’s argumentation‑scheme theory and Govier’s criteria for argument cogency. The protocol adapts to different reasoning frames, goes beyond multiple‑choice formats, and treats both the reasoning that precedes a verdict and its post‑hoc justification. We applied it to nine frontier models on 200 high‑ambiguity MoralChoice items, yielding 6,778 judge‑scored cells with 89.6% inter‑judge agreement on binary failure judgments. Models defended their reasoning well above the rubric minimum on every dimension. Failures clustered on grounds and sufficiency and correlated with epistemic hedging rather than argument length. Reasoning was defended better than post‑hoc justification across all models and Govier dimensions. The justification scheme often differed from the reasoning scheme on a substantial share of dilemmas (≥20% per model), even though value‑based practical reasoning dominated both tracks. The protocol catches strictly indefensible defences (self‑contradiction, false premises) and highlights difficulties in characterising the role of retraction in AI alignment, suggesting a need for more situated evaluations.

Review

Original Source: https://arxiv.org/abs/2609.05088

[h] Back to Home