NeFut Logo NeFut
中 Admin Login

[CS.AI] LUMOS: Tracing Parametric Knowledge from Training Data to Behavioral Outputs in LLMs

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#Machine Learning #LLM #Artificial Intelligence

Most existing analyses of large language models (LLMs) focus on the output side, drawing conclusions about what a model knows without confirming which training instances the model actually saw. This makes it hard to tell whether a correct answer stems from genuine generalization or simple memorization. To eliminate this ambiguity, we introduce the LUMOS diagnostic framework, which traces knowledge along the causal chain from training‑data exposure to behavioral output, leveraging the fully transparent training corpus of OLMo 2.

After verifying that the model has indeed encountered rare facts, we find that the internal representations separate these facts with high fidelity (84% separability) yet the models only express them behaviorally about 54% of the time. The retrieval gap narrows as model scale increases.

When the models are asked to self‑reflect on their own answers, they reliably self‑evaluate on trained content (83% accuracy) but drop to near‑random baseline performance (49%) on unseen content. Even with chain‑of‑thought prompting, confidence signals are inflated without improving calibration.

Overall, incorporating the training‑data axis into LLM evaluation turns speculative diagnoses into verifiable claims. We advocate that this axis become a standard component of knowledge assessment for LLMs.

Review: LUMOS offers a traceable experimental pathway to understand the source of LLM knowledge, exposing the gap between internal memorization and external behavior, and showing that self‑reflection and chain‑of‑thought are not universal calibration fixes.

Original Source: https://arxiv.org/abs/2610.02902

[h] Back to Home