When an informed adversary shares the audience of a constrained signalling channel, the signal that best protects the truth is the one that best describes it. On 108 confirmatory items we find that the adversary‑robust optimum aligns exactly with the salience pole identified in prior work.
Across a pool of 200,000 items the two differ on only 2,748 items—precisely where the previous salience‑to‑Bayes coordinate is undefined. Where defined, robustness is achieved by moving entirely from Bayesian discrimination to salience.
We introduce an adversary into a forced‑choice task (abstracted from Deception: Murder in Hong Kong). The adversary knows the target, observes the signal, and argues for the strongest wrong answer using a persuasion budget $\beta$. As $\beta$ grows, the optimal signal shifts from the posterior‑maximizing option to the margin‑maximizing one; at $\beta = 0$ the game reproduces the original oracle model with a listener temperature $\tau = 1$.
Key findings: 18.2% of the pool has an optimum that shifts under a finite budget, and each item’s critical budget is exact. Two adversary framings change the chosen option of seven language models on 30 to 77 of the 108 items, exceeding the exact no‑effect rate. Yet no measurement can determine whether this movement is toward the adversary‑aware optimum or toward salience, because the two options are identical. This is a structural limit, not a null result.
Practical advice: before evaluating adversary‑awareness, verify whether the robust target coincides with a heuristic target on the evaluation items. This diagnostic check is cheap.
Review