NeFut Logo NeFut
中 Admin Login

[CS.AI] What Makes a Terminal-Bench Task Hard? Distinguishing Genuine Hardness from Fake Hardness

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

Frontier benchmarks require tasks that current models cannot solve, yet a zero pass rate does not automatically imply intrinsic difficulty. The same all‑fail outcome may stem from a genuine capability gap, missing context, a broken reference solution, infrastructure failures, or a verifier that can be bypassed. In this work we investigate the issue using a frozen Terminal‑Bench 3 / Frontier‑Bench 0.1 production record comprising 1,081 pull requests, 639 scored tasks, 28,801 trials, and $105,933 of logged agent spend. We focus on the 125 tasks with no honest pass, aggregating task artifacts, reference‑solution runs, empty‑solution controls, adversarial trials, trajectories, telemetry, and review logs, then applying an ordered validity screen. Only 78 of the 125 survive as certified‑unsolved candidates. The remaining tasks include 14 with broken oracles, 8 dominated by infrastructure failures, 4 that are only passable via verifier bypasses, and 21 whose solvability is not certified by the available evidence. Hence, lack of saturation is not equivalent to genuine difficulty. The certified‑unsolved label is also narrow: it means the authored route passed, infrastructure did not dominate, no strict bypass was observed, and all evaluated agents failed, but it does not prove intrinsic hardness, verifier completeness, or failure at the intended capability. We further analyze rejected submissions and passing tasks to show that pass rate alone cannot explain why a task is difficult. Overall, our results suggest that frontier benchmarks should disclose the evidence behind all‑fail tasks before using them as capability claims.

Review

Original Source: https://arxiv.org/abs/2609.26826

[h] Back to Home