Abstract
All frontier large language models (LLMs) exhibit response drift—producing outputs that deviate from expert-validated references. However, the magnitude and structure of this drift remain uncharacterized by systematic human evaluation. We report a fully crossed evaluation in which 47 geographically diverse participants each assessed all 62 multidomain questions across ten frontier LLMs under blinded conditions, yielding 29,140 independent assessments.
Every model drifts, but drift magnitude varies substantially: eight models converge on a statistically indistinguishable ceiling (78-81% deviation), while two achieve lower deviation (47-49%). Drift profiles differ across six domains and 62 questions, with pairwise correlations among ceiling models exceeding r = 0.85. Automated similarity metrics explain less than 2% of variance in human judgments. These findings reveal that response drift is universal across frontier LLMs, domain- and question-dependent in structure, and accessible only through human-centered evaluation.
Blogger's Review: This study highlights potential issues in the practical application of current frontier LLMs, emphasizing the importance of human evaluation. Future model improvements should focus on understanding the roots of these drift phenomena to enhance model reliability.