Recent breakthroughs in AI are enabling scientists to make significant progress in mathematics, medicine, and materials science. New evaluation datasets are essential for further AI advancement, yet existing STEM datasets have notable shortcomings: model performance on them is saturated, leaving little room for meaningful evaluation; taxonomy distributions are skewed; most use multiple‑choice formats that do not reflect how researchers actually employ AI; and answers and rationales often suffer from inaccuracies due to contest‑driven collection and time‑constrained reviews. To address these gaps, we introduce Expert‑validated STEM QA, a high‑quality dataset created by 241 domain experts covering Physics, Chemistry, Biology, and Mathematics, with a total of 398 question‑answer pairs. Our construction pipeline follows four principled steps: (1) design a balanced taxonomy to ensure even coverage across topics; (2) apply a quality‑driven incentive scheme for contributors; (3) conduct multiple review rounds, revising items based on expert consensus; and (4) present the data in a verifiable Q&A format, making each answer and its rationale traceable. Empirical evaluation shows that state‑of‑the‑art frontier models achieve low performance on this benchmark, highlighting its difficulty and diagnostic value.
Review: By combining rigorous expert validation with a well‑balanced taxonomy, this dataset fills a critical void in STEM evaluation resources and offers a more realistic benchmark for future AI systems in scientific research.