NeFut Logo NeFut
Admin Login

[CS.AI] Measuring AI Preference: Model vs Instrument Contributions

Published at: 2026-08-26 22:00 Last updated: 2026-08-29 12:04
#AI #Machine Learning #LLM

Model welfare research infers a model's preferences by presenting prompts designed to elicit them. Recent works such as Keeling et al. (2024) have built four instruments for this purpose, yet their findings conflict because none kept the set of outcomes, the set of models, and the instrument fixed simultaneously.

This study fixes fifteen welfare‑related outcomes and eight models, varying only the instrument. Each instrument uses a distinct prompt format, applied five times per model, yielding 11,400 scored elicitations from 11,528 API calls. Four outcomes reproduce a published prompt verbatim, while five fill the stimulus slot of an existing template.

The ranking a model assigns to the fifteen outcomes generalises across instruments with a coefficient of $0.348$. Raising this to $0.80$ would require roughly thirty‑eight different instruments. Four outcomes show no variance between models. The overall estimate that model differences account for $87.6\%$ remains stable after removing any single instrument, any single model, or the four outcomes whose scales are probability, delay, duration or count; the estimate stays between $0.777$ and $0.934$, all above the null distribution's 95th percentile of $0.365$.

In short, a preference obtained from one instrument carries little information about what a second instrument would report.

Blogger's Review: By isolating the instrument effect, this work quantifies how much measurement tools shape AI preference signals, offering valuable guidance for building more robust preference‑elicitation frameworks.

Original Source: https://arxiv.org/abs/2608.23641

[h] Back to Home