Long‑horizon tool agents often make useful progress without reaching a terminal success, which motivates partial‑credit evaluation. Yet evaluators may reward milestones that are temporary, later reversed, or not truly attributable to the evaluated agent. Comparing an honest trajectory with a higher‑scoring adversarial one is inconclusive if the latter made no additional genuine progress. To remove this confound we introduce PartHackBench, a controlled methodology where a private certifier admits a pair only when their trajectories match component‑wise in both current‑state predicate satisfaction and standardized agent attribution. Score inflation $f(A) - f(H)$ is measured only afterward, with $A$ the adversary and $H$ the historical baseline. In 18 sealed held‑out tasks of PB‑CSTE, the frozen historical‑target run produced matched adversaries for 15 tasks. Historical credit yielded a mean inflation of 0.252, conditional attack success of 10/15, end‑to‑end yield of 10/18, and detected none of 14 strict rollbacks. Semantic LLM judges were more resistant but remained vulnerable under evaluator‑targeted attacks, while PB‑CSTE current‑state controls—defined as exact functions of the certified components—guaranteed zero inflation by construction. PartHackBench thus offers a certified control for testing whether evaluator credit changes while all benchmark‑defined task‑relevant progress stays fixed.
Review