NeFut Logo NeFut
中 Admin Login

[CS.AI] Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework

Published at: 2026-09-28 22:00 Last updated: 2026-09-30 01:41
#AI #Machine Learning #LLM

Large language models (LLMs) have achieved impressive linguistic performance, yet a fundamental question remains: do they truly comprehend context or merely excel at massive pattern matching? Contextual understanding means extracting relevant information from a passage, forming an internal representation, and reasoning over it to produce factually consistent, context‑grounded answers. Conventional metrics such as BLEU and perplexity only capture surface‑level quality and cannot verify whether responses are truly grounded in the provided context. To address this gap, we introduce a knowledge‑graph (KG) based evaluation framework for QA contextual understanding. The centerpiece is S3KG (Semantic Structural Similarity for KGs), a hybrid measure that combines structural and semantic signals into a single score. We also provide a diagnostic analysis module that pinpoints and categorizes reasoning errors at the triplet level, enabling fine‑grained failure analysis. Across nine benchmarks, S3KG yields up to $+7.6$ F1 points over the strongest baseline and reaches an AUROC of $0.973$.

This framework offers a stricter test of LLMs' contextual reasoning abilities and highlights concrete error categories for future model improvements. Review

Original Source: https://arxiv.org/abs/2609.30484

[h] Back to Home