This work examines item‑sensitivity in language models—whether a model’s choice depends on the specific input rather than its own prior output—and tests if this property truly indicates task competence. We employ a forced‑choice signalling task abstracted from the board game Deception: Murder in Hong Kong, where reference points (a fit‑maximising strategy, a posterior‑maximising strategy, and uniform random selection) are all computable in closed form.
Across seven language models, two model families, a post‑training ablation, and three independent scoring rules, all 21 model‑by‑rule cells exhibit reliable item‑sensitivity. Yet eight of those cells are statistically indistinguishable from a chooser that ignores the item and selects at random, and five score worse than random at describing the target. The correlation between item‑sensitivity and distance from random is only $r = 0.30$. We term this phenomenon consistency without alignment and argue it generalises to any evaluation that relies on item‑sensitivity, permutation consistency, or self‑consistency without an independent reference.
Additional findings show that a literal‑similarity baseline with no pragmatics outperforms most tested language models; adding a pragmatic layer over two similarity sources pushes choosers toward random rather than toward the Bayesian reference; and a standard labelled multiple‑choice format carries no measurable content signal in this setting. All results represent the model side of a pre‑registered instrument; a matched human condition has been designed and piloted but not yet collected.
Review