Large language models are increasingly deployed as cheap judges for output evaluation, data labeling, and system quality assessment. However, treating AI judgments as ground‑truth labels in formal statistical inference is inappropriate because AI can be biased or noisy, and hypothesis testing must strictly control type‑I and type‑II errors.
This work investigates how to conduct a valid hypothesis test under cost constraints by combining AI judgments with selective human verification. Consider a population of items with hidden binary labels. For a fixed pool of items the decision maker may: (1) query the AI directly; (2) send an item straight to a human; (3) after observing the AI report, escalate the item to human review; or (4) stop once enough evidence has accumulated.
We derive an information‑theoretic lower bound that captures the minimum cost required to meet prescribed error rates and introduce a report‑dependent information frontier that quantifies the value of AI information versus human verification. The lower bound can be expressed as $$C_{\min}= \frac{\log(1/\alpha)}{I_{\text{AI}}} + \frac{\log(1/\beta)}{I_{\text{human}}}$$.
Motivated by this characterization, we propose SCALE, a sequential cost‑aware policy that blends selective AI scoring with adaptive human escalation. SCALE is valid for finite sample sizes and matches the lower bound to first order as the target error probabilities vanish. To handle unknown AI output models, we estimate them using paired AI‑human pilot data.
Numerical experiments show that when either AI or human verification clearly dominates, SCALE behaves similarly to the corresponding AI‑only or human‑only test; when inexpensive AI judgments and selective human verification are both valuable, SCALE achieves the greatest cost savings.
Review