NeFut Logo NeFut
Admin Login

[CS.AI] BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification

Published at: 2026-08-26 22:00 Last updated: 2026-08-29 12:04
#AI #Machine Learning #LLM

We introduce BenchBench-Protocol, a benchmark for large language models that includes 149 protocol-modification tasks derived from scientists' real-world changes to published protocols. Adapting a published protocol to a new experiment is a routine task for wet‑lab scientists, and a correct modification must account for prior choices and downstream steps. Recent life‑science benchmarks have shifted toward open‑ended, rubric‑graded tasks, but these are usually authored by experts rather than reconstructed from actual modifications. BenchBench-Protocol tasks stem from the differences between a published protocol and a scientist's modified version, providing both the query and the weighted rubric elements for a correct response. The benchmark draws from 96 source protocols across nine wet‑lab biology domains and retains only tasks rated highly by domain experts. We evaluated nine closed‑source and open‑source models; Claude Opus 5 achieved the highest normalized rubric score at 59.2%, while other models ranged from 34.1% to 47.1%. Even when taking the best of ten attempts, the benchmark remains unsaturated. As models become increasingly helpful in life‑science research, assessing them on routine wet‑lab tasks becomes correspondingly important. BenchBench-Protocol offers a grounded assessment of wet‑lab reasoning and demonstrates the value of constructing benchmark tasks from real experiments.

Blogger's Review: This benchmark, built from authentic experimental modifications, enhances evaluation realism and difficulty, and is crucial for advancing the practical use of LLMs in laboratory settings.

Original Source: https://arxiv.org/abs/2608.23898

[h] Back to Home