NeFut Logo NeFut
中 Admin Login

[CS.AI] MolDesignBench: Evaluating LLM-based Agents for Scenario-grounded Molecular Design

Published at: 2026-09-24 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

Real‑world molecular design poses several challenges for large language model (LLM) agents: they must interpret design contexts, satisfy both explicit and implicit constraints, detect infeasible specifications, and reason over multi‑step tool outputs. Existing benchmarks largely focus on narrow, explicit constraints, only feasible problems, and single‑path solutions, which do not reflect practical demands.

To fill this gap we introduce MolDesignBench, a scenario‑grounded benchmark that more faithfully evaluates tool‑augmented LLM agents for molecular design. The benchmark comprises roughly 2K generation and optimization instances. Each instance blends implicit requirements embedded in a design narrative with explicit property and functional‑group constraints, and includes infeasible cases. Successful completion requires effective use of 17 specialized chemistry tools for tasks such as molecule construction, property prediction, and synthetic feasibility assessment.

Experiments across a range of frontier LLMs show low overall success rates, with the best model achieving only about $43\%$. Failures concentrate on implicit‑constraint reasoning, infeasibility detection, and tool‑reasoning. Fine‑grained failure‑mode analysis identifies implicit constraint interpretation and infeasibility detection as the primary bottlenecks.

MolDesignBench provides publicly available benchmark data, tool interface specifications, and evaluation code, establishing a rigorous testbed for future research on chemical agents.

Review

Original Source: https://arxiv.org/abs/2609.27349

[h] Back to Home