NeFut Logo NeFut
中 Admin Login

[CS.AI] Evaluating LLM-Generated Preference Distributions

Published at: 2026-10-03 22:00 Last updated: 2026-10-06 12:11
#Machine Learning #LLM #Artificial Intelligence

In this work we systematically examine the probability distributions produced by large language models (LLMs) for three choice domains: air travel, restaurant selection, and consumer products. Our experiments span nine open‑weight models and evaluate sampling under various temperatures, greedy decoding, and perturbations of prompts and item ordering. All models exhibit self‑coherence: the most probable outcomes stabilize quickly across repeated draws. At the same time, substantial discordance appears across model families and scales, with even the top‑ranked outcomes lacking consensus. This divergence persists across all three domains and remains robust to temperature changes, decoding strategies, and prompt tweaks. The findings indicate that the choice of model influences the generated preference distribution far more than the wording of the prompt, challenging the common assumption that sufficiently capable LLMs can replace survey respondents and yield similar distributions.

Review

Original Source: https://arxiv.org/abs/2610.01000

[h] Back to Home