To scale collective decision‑making, platforms such as Polis and Remesh enable online deliberation among thousands of participants. At this scale, users cannot review every opinion, resulting in extremely sparse voting data that misrepresent consensus, conflict, and minority support. Consequently, platforms increasingly rely on Preference Inference (PI) models to predict missing votes. This automation is not neutral: inferred preferences can artificially amplify, suppress, or reorder existing support patterns, reshaping how deliberation outcomes are interpreted. Moreover, we lack a systematic understanding of how current PI methods affect the overall preference landscape. To address this gap, we benchmark several PI approaches in real participatory‑democracy settings. Moving beyond conventional user‑centric evaluations that focus on individual prediction accuracy, we introduce a collective‑centric evaluation framework that measures whether inferred votes preserve salient properties of the broader preference landscape. We also provide the largest multilingual dataset of its kind, covering four consultations, over 90 k participants, 1 M votes, and 22 languages. Experiments show that models with comparable predictive accuracy can differ markedly in preserving collective structure, demonstrating that accuracy alone is insufficient for evaluating PI in democratic contexts. By offering this novel, collective‑centric benchmark, the work aims to support AI systems that scale deliberation without compromising the integrity of democratic outcomes.
Review: The study highlights the necessity of evaluating AI‑driven inference not just by per‑user metrics but by its impact on the whole preference ecosystem, safeguarding democratic deliberation from unintended algorithmic bias.