BioEVAL (BioEngineering Validation of AI and LLMs) is a benchmark created by 22 research groups worldwide to measure large language model reasoning on experimental tasks across bioengineering. The benchmark spans 11 major subfields plus uncategorized items, comprising 608 evaluation items: 380 multiple‑choice questions (359 retained after audit), 218 literature synthesis tasks, and 10 multimodal problems that require interpreting experimental images. Items were authored, reviewed by domain experts, and passed centralized quality control. After evaluation, a blinded cross‑group audit of the highest‑ and lowest‑accuracy MCQs identified 21 questions for revision or removal; these were excluded, and all reported MCQ results are based on the 359 retained items. We evaluated a range of cloud‑scale foundation and multimodal models (e.g., ChatGPT, Gemini, Grok) as well as locally deployable models that run on consumer‑grade GPUs. Models achieved up to 90% accuracy on MCQs, a similarity score of 0.72 on literature synthesis, and 80% accuracy on a small set of multimodal reasoning questions, with substantial variation across subfields. Leaderboard rankings highlight current capabilities, limitations, and development priorities for each task category. BioEVAL is maintained as an extensible benchmark with standardized protocols for ongoing expert item contribution and model evaluation.
Review