Human‑feedback alignment has turned language models into useful assistants and is often described as aligning them with humans. Yet the responses that people prefer from an AI need not be the responses they would themselves produce. We separate alignment with human preferences from alignment with human behavior, and show that preference alignment can make model behavior less human‑like even when both preferences and responses are entirely human‑derived. We call this phenomenon the Turing‑test gap. Preference alignment preserves the human response distribution only under a restrictive condition, and empirical evidence shows that real human preferences do not consistently satisfy it. Experiments reveal that increasing the weight of preferences reduces the likelihood of human‑like responses, regardless of direction, and the same gap appears under standard DPO. These findings establish human‑likeness as an explicit dimension of alignment rather than an automatic consequence of preference alignment.
Review