Repeated‑sampling evaluations often extrapolate pass@k far beyond the number $n$ of samples collected per problem. In the pooled/random‑task conditional‑Binomial model we show that a fixed‑$n$ success count identifies only the $n$ free moments of the latent per‑task success distribution. Consequently, direct pass@k is identifiable only for $k \le n$, even with arbitrarily many exchangeable tasks under the same rollout budget. This is stronger than the usual observation that the estimator is undefined beyond $n$, because it characterizes the missing information in a fixed‑depth count‑law experiment.
We provide exact count‑law‑preserving constructions that yield incompatible extrapolations and point out the exceptional case of a unique extension. Sharp population‑identified intervals are computed via Hausdorff principal representations.
On the public 10,000‑rollout‑per‑problem release of Brown et al., a counterfactual evaluation with $n=16$ leaves the failure rate at $k=1000$ ambiguous by factors ranging from 1.5 to over 2,600 across four MATH/GSM8K/CodeContests configurations. Calibration shows that intermediate‑scale failure share alone does not determine the interval width.
Our result does not reject parametric inference‑time scaling laws; it supplies a non‑parametric baseline against which their assumptions can be evaluated. We also present an exact, conservative one‑coordinate finite‑task confidence certificate and a reporting standard that separates direct estimates, identified sets, and model‑conditioned forecasts.
Review