Embodied agents promise automation of scientific experiments, yet their progress is hampered by the lack of reliable, systematic evaluation environments. Existing simulation‑based laboratory benchmarks rely heavily on manual task engineering, making it difficult to compile diverse scientific protocols into executable, verifiable embodied tasks at scale.
SciHorizon‑eLab addresses this gap by treating scientific embodied task construction as a compilation problem. Given a natural‑language protocol, the system proceeds through three stages: semantic grounding maps protocol concepts to simulation entities; executable task synthesis produces manipulation programs; multi‑stage simulation certification checks semantic preservation and emits step‑level success specifications.
The pipeline yields semantically aligned environments, runnable manipulation scripts, and per‑step success criteria, while also generating reproducible expert demonstrations and execution traces. Using this workflow we built \BenchName, a ready‑to‑use benchmark of 300 certified tasks covering a wide range of laboratory operations. The benchmark supports hardware‑in‑the‑loop (HIL) execution, reproducible expert‑demo generation, and ordered step‑level evaluation.
Across representative tasks the strongest policy achieves an average success rate of only 49.7%, and further analysis reveals pronounced weaknesses in human‑agent and agent‑agent coordination. The code, benchmark data, and evaluation toolkit are publicly released at https://github.com/SciHorizon-elab/SciHorizon-elab.
Review: SciHorizon‑eLab offers a systematic route from natural‑language protocols to executable embodied tasks, establishing a scalable benchmark for scientific automation. However, current policies struggle to surpass a 50% success threshold, indicating substantial challenges remain in fine‑grained control and collaborative coordination.