NeFut Logo NeFut
中 Admin Login

[CS.AI] Verifiable, Articulable, and Tacit Components of Preference

Published at: 2026-10-06 22:00 Last updated: 2026-10-08 01:25
#AI #Machine Learning #LLM

We introduce a large‑scale labeled preference dataset CreativePreferences, comprising 2.8 M texts across seven creative domains, annotated with 317 M human preference judgments and organized into 42 benchmark tasks. Labels are modeled using executable programs, rubric banks, and densely trained models denoted V, A, and VAT respectively. Experiments reveal two systematic gaps: an articulability gap (VAT‑VA) and a verifiability gap (VAT‑V). A novel measurement approach—based on capture‑recapture—discovers articulable and verifiable metrics, filters spurious variables, and estimates the value of undiscovered metrics, yielding upper and lower bounds for each gap. These gaps appear in every domain, even those traditionally deemed fully verifiable such as mathematics and software engineering, as well as claim‑ and novelty‑centric domains like news, patents, and peer review. Gap magnitude varies by domain; peer review and creative writing exhibit the largest articulability gaps, and the gaps widen as more judges participate, consistent with Collins’s theory of collective tacit knowledge. We demonstrate two key consequences: (1) for human‑generated tasks, the full model (VAT) aligns more closely with human preferences, often contradicting explicit criteria; (2) analogous to Goodhart’s law, articulating preferences shifts them away from their tacit dimension. Based on these findings we recommend when prompting is appropriate, how learning mechanisms might be improved, and when judgments should remain with humans.

Review

Original Source: https://arxiv.org/abs/2610.03025

[h] Back to Home