We investigate whether domain‑specific fine‑tuning remains beneficial for open‑ended scientific reasoning in astronomy language models. To this end we built a curated QA benchmark from publicly available 2017‑2026 Olympiad‑style materials, comprising 300 free‑response questions: 204 text‑only and 96 linked with images.
We evaluated open‑weight models, API‑served general‑purpose multimodal models, and astronomy‑specialized models. Evaluation relied on judge‑based correctness and several reference metrics. Strong general‑purpose models achieved the highest overall correctness baseline, yet analyses of metric agreement, judge sensitivity, benchmark composition, and modality revealed variations that a single leaderboard cannot capture.
The results suggest that domain specialization should be treated as a task‑ and deployment‑dependent property rather than a universal improvement. Domain‑specific evaluation is crucial for selecting appropriate models, capabilities, and evaluation criteria in scientific workflows.
Review: Generalist models now rival specialized ones in raw performance, but careful alignment with task requirements and evaluation protocols remains essential for scientific reasoning.