We evaluate Scientific Agents, an open‑source collection of 503 profession‑specific system prompts (AGENTS.md), using Gemini 3.8 Flash via OpenRouter in the Pi agent harness. Four controls are compared: a minimal baseline ("You are a helpful assistant"), the opening role sentence of the profile, a generic scientific rigor guide, and a profile from an unrelated domain. The study spans nine text‑based science benchmarks (4,531 sampled questions, 100 matched profiles); after API‑error retries, 4,488 items completed all five conditions and were scored with automated rule‑based grading. The average accuracy difference between profile and baseline is –0.6 percentage points (95 % bootstrap interval [‑1.5, +0.2]), showing no clear gain. Profiles generate 1.5‑2.3 × more output tokens and cost 2.2‑4.5 × more per successful call. On 60 tool‑using BioMysteryBench bioinformatics problems (three runs per condition), baseline solve rate is 56.7 % versus 46.7 % for the profile, a –10 % gap driven by more frequent token‑ or time‑limit stops under the profile. Longer prompts unexpectedly help on SuperGPQA: the short baseline yields a correct first‑pass answer on only 54.0 % of items, versus 71.6 % with the profile, likely due to prompt length or formatting rather than domain expertise. The conclusion is that, for the tested model and tasks, loading full profession profiles by default does not improve accuracy and raises cost substantially; whether selective retrieval of profile sections or open‑ended scientific tasks would change this remains to be tested.
Review