NeFut Logo NeFut
Admin Login

[CS.AI] NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment

Published at: 2026-09-12 22:00 Last updated: 2026-09-15 01:15
#AI #Machine Learning #LLM

NovGauge is a fine-grained benchmark for assessing the novelty judgment capability of large language models (LLMs). The dataset draws from ICLR reviewer overlap claims and survey co‑citations, comprising 619 paper pairs and 50 multi‑paper sets. Each instance is independently labeled along three dimensions—task, problem, method—to capture application goals, technical challenges, and solution approaches. A cascading diagnostic pipeline first verifies per‑dimension correctness, then checks evidence grounding, and finally evaluates logical support. Evaluation of 18 LLMs reveals hallucination rates ranging from 0% to 39% across dimensions; among non‑hallucinated correct positives, over 70% of cited evidence fails to logically back the stated reason. The top model, GPT‑5.5, attains Verified F1 scores between 43% and 72% across dimensions, while most models lose more than half of their raw F1 after faithfulness verification. These findings indicate that current LLMs remain far from reliable scientific novelty assessment, especially when correctness depends on faithful evidence grounding.

Review

Original Source: https://arxiv.org/abs/2609.11234

[h] Back to Home