NeFut Logo NeFut
Admin Login

[CS.AI] Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools

Published at: 2026-09-10 22:00 Last updated: 2026-09-12 06:35
#AI #Machine Learning #LLM

We introduce the ABLE benchmark to assess how large language model (LLM) agents employ biological AI models (BAIMs) such as ProteinMPNN and AlphaFold3 within dual‑use protein design pipelines.

ABLE comprises tasks across three axes—structure retrieval, sequence generation, and design validation—to measure performance in information gathering, tool selection, and actual tool execution.

Fifteen state‑of‑the‑art models were evaluated; seven refused all tasks, while the remaining models showed marked variability. Claude Sonnet 4 and Gemini 3 Pro achieved the highest scores across retrieval, selection, and usage sub‑tasks.

When a subset of tasks was compared against an expert human baseline, results indicate that current LLMs can substantially lower the barrier to protein design, yet they remain inconsistent in planning, strategy formulation, and integrating biological knowledge with tool operation.

Review

Original Source: https://arxiv.org/abs/2609.05818

[h] Back to Home