Large language model (LLM) based coding agents have made notable progress on repository‑level software engineering tasks. Existing repository benchmarks usually start from a human‑identified issue and only check whether a patch satisfies a functional signal. We introduce SWE‑Prometheus, a benchmark targeting the broader task of improving repository engineering governance. Each task supplies a fixed snapshot and an open‑ended objective, requiring the agent to identify risks, prioritize interventions, and verify the resulting changes.
SWE‑Prometheus evaluates six governance dimensions using paired evidence, clean‑environment probes, behavior gates, and two independent teacher ratings of the same evidence. The benchmark comprises 60 repositories; ten models are evaluated on a shared public subset of 22 repositories, where the mean Normalized Governance Improvement (NGI) ranges from 0.0568 to 0.5760 and observed behavior‑breakage rates range from 0% to 23%.
On a frozen batch of ten repositories, a repository‑blind template achieves a mean NGI of 0.272, but its gains are concentrated in Tests & CI, Quality Gates, and Documentation, with no improvement in Reproducible Environment or Dependency & Security. This baseline distinguishes between merely adding governance artifacts and producing execution‑backed improvements.
The no‑op condition has a median NGI of zero and a standard deviation of 0.073; two teachers agree exactly on 57 of the 60 dimension scores for the same no‑op evidence. For the two highest conditional‑mean systems, common‑valid NGI is similar, while full‑pool comparisons that include behavior failures favor Kimi‑K3.
These findings demonstrate that repository‑governance evaluation should report improvement, behavior preservation, evidence quality, and coverage together.
Review