NeFut Logo NeFut
Admin Login

[CS.AI] Moral Competence Before Moral Content: Why LLM Agents Lack Prerequisites for Coherent Alignment

Published at: 2026-09-08 22:00 Last updated: 2026-09-09 09:08
#algorithm #AI #Machine Learning

AI alignment demands that systems follow human norms, values, or intentions. Under value pluralism there is no single correct target, but a shared prerequisite is that system behavior expresses a coherent policy: a mapping from situations to verdicts that remains invariant when morally relevant features are preserved and changes when they do not. We introduce four structural conditions for such policies—verdict stability, monotonicity, decisiveness, and Pareto viability. Together they measure a form of moral competence that can be evaluated from behavior alone, without reference to external moral standards or expert baselines, thus providing a structural floor for alignment rather than a normative target.

We demonstrate this methodology across three simulated deployments where LLM‑based agents face moral dilemmas. The study evaluates nine frontier models using a factorial design of five paraphrases, five escalation levels, and three dominance conditions. No model exhibits a coherent policy across the deployments: surface‑form perturbations alone cause verdict‑rate shifts of up to $99$ percentage points at a single escalation level, and success in one scenario does not predict competence in another.

These results suggest that current LLM‑based agents are not the kind of objects to which alignment can meaningfully apply, as they lack the necessary moral competence.

Review

Original Source: https://arxiv.org/abs/2609.05036

[h] Back to Home