NeFut Logo NeFut
Admin Login

[CS.AI] TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

Published at: 2026-09-08 22:00 Last updated: 2026-09-09 09:08
#AI #Machine Learning #LLM

Autonomous coding agents are being proposed as AI‑scientist systems that can run analyses and write research reports, yet carrying out a prescribed analysis is not the same as making a discovery. Existing benchmarks are built around a hidden target study, with tasks, data and rubrics that reward reproducing its result. TruthInsightBench is configured for discovery instead. It offers 40 blind tasks drawn from 40 peer‑reviewed papers across 10 scientific domains, exposing only a neutral scientific objective and frozen data while withholding source conclusions, expected values and analysis paths, leaving the agent to infer what claim the data support. Scoring is performed by a fixed LLM‑based judge that rates the evidentiary maturity of the agent’s claims along six dimensions, operationalized as 29 artifact‑grounded items, with automated deterministic aggregation and no per‑instance human grading, enabling fully repeatable evaluation as agents evolve. On a single frozen base model, four coding agents form a narrow plateau (58.4‑60.3 out of 100) with no statistically reliable pairwise separation: they execute and document analyses competently, showing relatively strong evidence auditability and novelty, but largely lack the discriminating acts that establish a trustworthy claim—controls, robustness checks, falsifiability and cross‑dataset generalization. The bottleneck is scientific judgment rather than coding, and genuine discovery remains out of reach. TruthInsightBench turns this gap into a measurable target; data and scoring code are available at https://github.com/TruthInsight-stack/TruthInsightBench.

Review

Original Source: https://arxiv.org/abs/2609.05079

[h] Back to Home