NeFut Logo NeFut
Admin Login

[CS.AI] QVAC Genesis III: A Large-Scale High-Quality Open Synthetic STEM Corpus for Efficient LM Pre-Training

Published at: 2026-09-18 22:00 Last updated: 2026-09-20 12:54
#AI #LLM #Open Source

High‑quality pre‑training data is a major bottleneck for education‑focused and STEM‑specific language models targeting edge AI and on‑device deployment, where token budgets are extremely tight. While large organisations keep scaling models on private corpora, the open ecosystem lacks synthetic STEM datasets that provide high per‑token learning value for small models. To close this gap we introduce QVAC Genesis III, a 191.43 B‑token synthetic corpus covering 19 domains, multiple difficulty levels and educational styles. It is built with a dual‑generation strategy: a weak edge‑scale student model is used for targeted teacher distillation, converting its failures into corrective explanations and expanding its successes into contrastive, option‑level reasoning over all answer choices. We also propose an LLM‑as‑a‑parser evaluation protocol that extracts final answers from free‑form outputs while tracking both accuracy and answer validity. Controlled from‑scratch ablations with 1.7 B‑parameter models show that models trained on QVAC Genesis III consistently outperform those trained on the open‑source synthetic corpus Cosmopedia‑v2 and the publicly released Cosmo‑1B model across ARC, GPQA Diamond and MMLU STEM benchmarks, achieving up to +28.57 % on ARC‑E, +21.35 % on ARC‑C and a Valid Answer Rate of up to 99.45%.

Review

Original Source: https://arxiv.org/abs/2609.19513

[h] Back to Home