Large language models (LLMs) can serve as judges to enable scalable evaluation, yet their judgments are often sensitive to answer order and may still diverge systematically from human preferences after removing order effects. DIAL introduces a unified framework that combines abundant LLM comparisons with limited human comparisons to separate judge‑specific order effects, learn a shared structure of position‑debiased LLM preferences, and adaptively calibrate this structure toward the human‑preference target. Theoretically, DIAL studies three aspects: (i) identifiability of latent LLM preferences, order effects, and human calibration; (ii) adaptive estimation that balances LLM anchoring against scarce human evidence; (iii) fixed‑weight uncertainty quantification for the calibrated human preference. Empirically, the authors evaluate debiasing and human alignment separately in controlled simulations and on three human‑preference benchmarks, showing that DIAL remains robust to unbalanced response order, achieves strong human‑aligned rankings with limited labels, and adapts toward human evidence when LLM information is imperfect. A real‑data study collected over 410K judgments from 21 LLM judges in both display orders, providing a valuable resource for future work on LLM‑judge bias, heterogeneity, and human alignment.
Review: DIAL demonstrates a practical route to achieve efficient human alignment under limited annotation by leveraging structured debiasing and adaptive calibration, offering new insights for fairness and reliability in automated evaluation systems.