NeFut Logo NeFut
Admin Login

[CS.AI] Stress Testing LLM Agents in Robotic Chemistry Labs

Published at: 2026-07-29 22:00 Last updated: 2026-07-30 03:24
#algorithm #AI #Open Source

AI is typically evaluated through knowledge, reasoning, and plan generation; however, scientific agency requires reliable physical actions and adaptability to evidence. This study employs a robotic chemistry laboratory as a physical-world testbed to measure scientific agency.

A total of 4,608 trials were conducted using 45 modular workstations exposed as machine-readable skills. Only 3.3% of trials produced expert-assessed executable workflows under laboratory constraints, with the best system achieving just 28.1%.

Long-horizon planning was particularly challenging, as only three executable workflows exceeded 30 operations, with the longest containing 44 operations. Throughout five rounds, experimental feedback led to local adjustments, but there was no workflow-level replanning or redesign of analytical methods.

By quantifying physical executability and evidence-driven replanning, this study provides an evidence-based assessment of deployment readiness and a diagnostic framework to guide closed-loop improvements for physically grounded autonomous research.

Blogger's Review: This research offers critical empirical data on the application of large language models in physical experiments, highlighting current limitations in handling complex tasks, particularly in long-term planning. Future research should focus on enhancing model adaptability in dynamic environments.

Original Source: https://arxiv.org/abs/2607.23045

[h] Back to Home