NeFut Logo NeFut
中 Admin Login

[CS.AI] MOF-VERIFY: A Failure-Aware Agentic Harness for MOF Hypothesis Verification

Published at: 2026-10-06 22:00 Last updated: 2026-10-08 01:25
#AI #Machine Learning #Artificial Intelligence

Large language models are increasingly integrated into AI‑driven materials co‑science, yet the trustworthiness of the verification pipeline is still uncertain. Metal‑organic frameworks (MOFs) pose a tough challenge because a single structure can appear under multiple identifiers, synthesis outcomes depend heavily on experimental conditions, evidence is scattered across heterogeneous sources, and some hypotheses require computational validation beyond literature. We introduce a diagnostic benchmark covering four task families: structural grounding, synthesis‑condition verification, evidence‑sufficiency checking, and ML‑based computational verification. For tasks T‑MOF‑1‑3 we evaluate under closed‑book, retrieval‑enabled, and oracle‑evidence settings to pinpoint failures in knowledge access, evidence acquisition, and reasoning; T‑MOF‑4 focuses on computational verification. Guided by these diagnosed failure modes, we develop MOF‑Verify, a failure‑aware agentic harness that addresses structural, literature, evidence‑sufficiency, and computational bottlenecks before issuing a final verdict. Across multiple backbone LLMs, MOF‑Verify substantially outperforms direct inference and retrieval baselines in hypothesis‑verification accuracy. The benchmark datasets are released at https://github.com/IMMS-Ewha/MOF-Verify-Benchmark.

Review

Original Source: https://arxiv.org/abs/2610.03056

[h] Back to Home