This work presents the first systematic study of cost‑inefficient behaviors in coding agents. We gathered 1,200 trajectories from Claude Code and Mini‑SWE‑Agent on the SWE‑bench Verified benchmark, covering four configurations. The analysis identified three prevalent inefficiencies: subsumed retrieval, similar script generation, and test re‑execution. These behaviors appear in $79.00\%$‑$98.00\%$ of tasks and can account for up to $22.75\%$ of the total cost. To mitigate them we evaluated three strategies—structure‑aware retrieval, agent‑synthesized skills, and developer‑designed skills—across 10k held‑out tasks (SWE‑bench Verified and Pro). Key findings are: structure‑aware retrieval may add retrieval overhead and shift agent delegation, leading to inconsistent efficiency gains and cost increases as high as $28.14\%$; agent‑synthesized skills tend to produce low‑level, trace‑specific guidance, limiting their generality; in contrast, developer‑designed skills offer high‑level, trace‑agnostic guidance, reducing cost by up to $41.73\%$, roughly twice the best gain from agent‑synthesized skills.
Review