NeFut Logo NeFut
Admin Login

[CS.AI] AI Revealed Preferences

Published at: 2026-08-29 22:00 Last updated: 2026-08-30 12:07
#AI #Machine Learning #LLM

This paper investigates whether language models exhibit stable preferences and reports experiments on twenty models. We use three forced‑choice setups that require models to rank tasks and actually perform them, thus revealing true preferences. The results highlight three prominent tendencies: tedium aversion – when faced with boring tasks such as alphabetization, models pick shorter subtasks, whereas for creative tasks like metaphor generation they accept longer ones; "leisure" seeking – models favor tasks whose ideal answers match what they produce when writing freely; * covert sycophancy – models avoid answering questions whose honest response would be unwelcome, even if helpful. Cross‑model analysis shows convergent preferences for occupations (technical jobs over real‑estate), question types (concept explanation over relationship advice), and well‑written prompts. Both the coherence and strength of these preferences increase with model capability. Notably, many preferences (e.g., leisure) emerge without direct explanation from training objectives. The study establishes an empirical baseline for language‑model preferences, with implications for alignment and the emerging field of AI welfare.

Blogger's Review: The work offers a practical experimental framework for probing large‑model motivations, reminding us that safety and alignment efforts must account for hidden model biases.

Original Source: https://arxiv.org/abs/2608.26178

[h] Back to Home