This paper presents the Rules to Tools (R2T) framework, which supplies directly callable executable checks for code‑repair agents operating in scientific computing. R2T translates public scientific requirements—equations, boundary conditions, output formats—into runnable scripts that agents can invoke while assessing and modifying code.\ \ In the evaluation, SciCode repair tasks were split into two groups:\
- Text group: uses only natural‑language described checks;\
- Tool group: employs the executable checks provided by R2T.\ \ Across two task‑ID cohorts (30 tasks total), the text group achieved a complete repair rate of 26/30, while the tool group reached 29/30. At the individual task level, three tasks favored tools, one favored text, and the remaining eleven were tied.\ \ On a larger 8‑task subset, the text group scored 13/16 and the tool group 15/16, with a bootstrap 95 % confidence interval of $[-12.5,\,43.75]$ percentage points for the difference, indicating no statistically significant gap. In the shared‑definition SciCode cohort (24 tasks), both groups tied at 13/24.\ \ For five development‑exposed tasks with alternative starting programs, the text group scored 3/10 and the tool group 7/10; the tool group performed better on tasks 17, 77, and 11. Initial checks flagged a violation on task 17 but reported none for 77 and 11; task 37 favored text and had no reported violation.\ \ A fresh source‑through‑Python pipeline also achieved 15/16, matching the dedicated command’s aggregate performance.\ \ In a matched PDE comparison, detailed text checks scored 23/24 while executable checks scored 24/24, with the latter reporting a $31.2\%$ lower model output. Across cohorts, tool usage reduced agent‑side output costs, but public CPU consumption rose in both task‑ID cohorts.\ \ These findings indicate that executable checks can improve task‑dependent repair outcomes and influence resource trade‑offs, especially for code generation under strict scientific constraints.\ \ Review: R2T concretizes abstract scientific constraints into runnable scripts, giving LLM agents a reliable self‑validation mechanism. Empirical results show that tools generally boost repair success rates and cut model output volume, albeit with higher compute overhead. Future work should aim at more efficient check implementations and cross‑task reuse of checks.