ScientistTwo is a fully autonomous multi‑agent framework designed to realize problem‑driven AI research. Given a fundamental challenge posed by a human expert, the system first establishes state‑of‑the‑art baselines on public datasets, then formulates novel hypotheses and dispatches specialized agents to handle theory derivation, experiment design, and code implementation, closing the loop from problem definition to paper drafting.
During the experimental phase the framework automatically selects diverse datasets and metrics, runs large‑scale experiments, and refines methods through automated ablation studies. Afterward, a simulated peer‑review rebuttal engine conducts self‑validation to ensure reproducibility and academic rigor.
Benchmarking on papers accepted at top conferences such as ICLR, ICML, and NeurIPS shows that ScientistTwo can independently generate expert‑level manuscripts and executable codebases, outperforming current human state‑of‑the‑art models and achieving higher average scores under automated AI reviewers. These results demonstrate that ScientistTwo has moved beyond a mere assistive tool to become an autonomous scientific pioneer capable of expanding the frontier of human knowledge.
Review