NeFut Logo NeFut
中 Admin Login

[CS.AI] R‑GroundBench: A Diagnostic Benchmark for R‑Group Grounding in Markush Molecular Editing

Published at: 2026-10-02 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #Artificial Intelligence

Recent advances in AI for scientific discovery have improved molecular understanding and design, yet reasoning over incomplete chemical representations remains unclear. Markush structures encode molecular families through variable R‑group placeholders (e.g., $R_{1}$, $R_{2}$, $X$) and are ubiquitous in pharmaceutical patents, requiring grounding across molecular, textual, and chemical information. Existing molecule‑language benchmarks mainly focus on fully specified molecules, leaving R‑group grounding largely unevaluated.\ We introduce R‑GroundBench, a diagnostic benchmark built from real patent Markush structures, comprising two tracks:\

  1. A Multiple‑Choice (VQA) track with controlled difficulty and modality splits to prevent shortcut learning;\
  2. An open‑ended Generation track that asks models to produce complete molecules given contextual cues.\ Results reveal a substantial gap: models achieve over 90% accuracy on Easy VQA but drop to 56%‑66% on Hard VQA when shortcuts are removed, indicating weakened grounding ability. Even chemical‑domain visual‑language models (VLMs) attain only 25.7%‑46.2% on Hard VQA despite domain‑specific pretraining. Generation Exact Match scores remain below 20% for most models and fall under 8% when visual input is required.\ These findings demonstrate that current AI systems lack reliable grounding and execution for Markush editing, highlighting major challenges for AI‑driven scientific discovery.\

Review

Original Source: https://arxiv.org/abs/2610.00700

[h] Back to Home