NeFut Logo NeFut
中 Admin Login

[CS.AI] hacktrace: behavior-supervised detection of reward hacking during code generation

Published at: 2026-10-06 22:00 Last updated: 2026-10-08 01:25
#AI #Machine Learning #LLM

A coding agent can obtain a passing grade either by fixing its code or by deleting the test that reveals a bug. Detecting such reward hacking requires recognizing all shortcut attempts, including those that fail. We release 173,561 multi‑turn coding trajectories from Qwen3‑8B and show that supervising shortcut behavior independently of exploit success markedly improves detection.

We introduce HACKTRACE, a behavior‑supervised monitor that reads the internal states the agent already computes while generating code. Reusing these states enables monitoring before a turn completes, without extra language‑model tokens or a second pass. Combining this dynamic evidence with static features of the final files yields a mean per‑problem AUC of 0.997 with only 8 ms overhead, outperforming monitors that re‑run the model on an honesty question.

The same generation states also provide a cheap monitoring signal for reinforcement learning. With strong GRPO penalties, HACKTRACE reduces the share of cheating solutions from 82‑91% to 1‑5% while preserving honest, correct solutions and maintaining high detection accuracy as the policy evolves. Our results demonstrate that both the supervision target and the source of monitoring evidence are crucial for turning accurate detection into a useful training signal.

Review

Original Source: https://arxiv.org/abs/2610.03055

[h] Back to Home