NeFut Logo NeFut
中 Admin Login

[CS.AI] CompMat-Bench: Benchmarking AI Agents for Computational Materials Science

Published at: 2026-10-02 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

CompMat-Bench is a benchmark for evaluating AI agents in computational materials science, comprising 94 tasks extracted from recently published simulation studies. Each task corresponds to a single step in a research workflow, requiring the agent to prepare inputs for costly simulations and to analyse the resulting outputs. All simulations are reproduced beforehand, providing ground‑truth inputs and results so that evaluation can avoid rerunning expensive calculations and can use fixed grading rules without an LLM judge.

The benchmark defines four evaluation settings: single‑task and multi‑task workflows, each with either full methodological guidance or reduced guidance. With full guidance, agents built on three leading LLMs achieve pass rates between 66.0% and 90.4% across the 94 tasks. Longer workflows or reduced guidance lower performance, but in different ways: weaker agents suffer under both conditions, while the strongest agent only drops noticeably when a long workflow is combined with reduced guidance.

Failure analysis attributes most errors to scientific mistakes (e.g., misinterpreting material properties or choosing inappropriate simulation parameters) rather than to software‑usage errors. CompMat-Bench thus offers a common platform for comparing agents on authentic materials‑research steps and for probing their failure modes.

Review

Original Source: https://arxiv.org/abs/2610.00636

[h] Back to Home