Large‑scale text annotation typically relies on a codebook crafted by experts, guiding AI annotators to produce consistent labels. Traditionally, refining such a codebook takes months because experts must review massive annotation outputs and iteratively adjust rules. This work proposes using disagreement among large language models (LLMs) to pinpoint the samples that most urgently need expert attention, compressing a multi‑month process into a few days.
We explore three expert‑feedback modalities driven by cross‑model disagreement:
- Codebook Verifying – when multiple LLMs assign different labels to the same text, the system generates revision suggestions that experts edit directly;
- Question Answering – experts answer targeted questions about the conflicting cases, helping models understand the source of disagreement;
- Rationale Labeling – experts provide reasoning for their chosen label on disagreement cases, allowing the model to learn more precise decision criteria.
Experiments on thousands of tutoring‑session transcripts show that Rationale Labeling yields the highest LLM labeling accuracy (64.9%), surpassing the expert‑revised codebook baseline (57.8%). The best Question Answering configuration also outperforms the baseline with 60.5% accuracy.
The core workflow is:
- Run several LLMs on the same dataset;
- Compute label distribution and select samples with the highest disagreement score;
- Present these samples to experts for focused feedback;
- Feed the expert feedback back into the LLMs to update the codebook or fine‑tune the models.
This loop enables rapid convergence to a high‑quality annotation scheme with minimal human effort.
Review