Leveraging frontier models such as Claude Opus as meta‑agents to synthesize terminal tasks and verifiers is becoming common, yet a runnable Docker image and test suite do not guarantee a faithful end‑to‑end training pipeline. We introduce a meta‑agent pipeline that isolates three failure modes: benchmark invalidity, harness brittleness, and reward misalignment. Prompt redesign and context extension boost baseline solvability by 5.6×, but a 9B model caps at 81.3% mean pass@2 within 20 steps on Claude Opus‑generated tasks. Adding hard tasks without altering the training configuration drops mean pass@2 to 20.6%, providing strong evidence that the solvability band is model‑specific. These results suggest that meta‑agent reliability requires solvability‑band calibration, verifier audits, and infrastructure error accounting as primary evaluation criteria rather than post‑hoc diagnostics.
Review