Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty. We model a fixed grid of $M$ tasks with $L$ binary paths per task under a hard budget $(M+t)K$, assuming each path costs at most $K$ responses or episodes. For fixed $L \ge 3$ and $0<\alpha \le 1/12$, the optimal expected interval width on the worst pure cohort is $\Theta_{\alpha,L}\big([M(t+1)]^{-1/2}\big)$ when every task is observed, and $\Theta_{\alpha,L}\big([M(t+\sqrt{M})]^{-1/2}\big)$ when omission is allowed. The lower bounds cover adaptive hard‑budget policies, while fixed random‑subset designs attain both rates via disagreement certificates. A joint mean/disagreement interval turns the task‑covering law into a practical finite‑budget inference tool. In an equal‑budget LiveCodeBench replay with 16 models, 880 tasks, and five outputs per task, the task‑covering design reduces median point‑estimation MSE by 87.0% relative to pooled uniform sampling, and the joint certificate yields narrower confidence intervals in 15/16 panels, cutting median interval width by 30.6%. Finite‑regime analysis shows that at the evaluated scale task coverage is the effective choice and clarifies how cohort size and within‑task agreement determine the useful operating region. Together, the sharp laws and fixed‑budget evidence make replication and task coverage explicit design variables for information‑efficient repeated evaluation.
Review