NeFut Logo NeFut
中 Admin Login

[CS.AI] Sequential Knowledge Editing Undermines Evidence Discrimination Without Hurting Accuracy

Published at: 2026-09-26 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

This paper investigates an overlooked side‑effect of sequential knowledge editing on large language models: the degradation of evidence discrimination while standard accuracy remains intact. Conventional editing benchmarks check three aspects—whether the edited fact changes, whether paraphrases stay consistent, and whether unrelated answers remain unchanged. A model can pass all three yet lose the ability to judge which retrieved documents are trustworthy for facts that were never edited.

We introduce an arbitration metric: for a fixed query, passage, and two candidate strings, we compare the log‑odds the model assigns to its remembered answer versus the answer asserted by the injected passage, before and after editing.

Using a conservatively tuned LoRA, we performed 1,000 sequential edits on Qwen2.5-7B‑Instruct. MMLU scores stayed unchanged to four decimal places, but the spread of the arbitration quantity across untouched facts dropped by 36%. Selective prediction deteriorated, the area under the risk‑coverage curve rose by 0.107 (versus 0.005 for a norm‑matched perturbation with the same MMLU drop), and the error on the model’s most confident quarter of arbitration decisions increased from 0.217 to 0.342. This is not a loss of overall capability. Random perturbations of five severities that reduced MMLU from 0.6275 to 0.3725 caused less harm to the arbitration metric (0.088) than MEMIT did at 0.6050 (0.102).

The effect persisted across three random seeds, two model families, two datasets, two probe‑disjointness criteria, three prompt templates, and paraphrased queries. Layer ablation on saved weight deltas showed the impact is distributed: no single layer reproduces it, and removing any layer recovers roughly half of the degradation. With a frozen retriever, retrieval accuracy fell from 0.592 to 0.46.

A secondary, perhaps more practical finding is that three of five model‑method pairings collapsed to chance‑level MMLU after 1,000 sequential edits under published hyper‑parameters, even though edit success remained at 1.00 and locality metrics appeared clean. Sequential‑editing evaluations that never measure capability would miss this failure mode.

Review: Sequential knowledge editing can preserve traditional accuracy metrics yet severely impair a model’s ability to discern reliable evidence, highlighting the need for capability‑aware evaluation in future research.

Original Source: https://arxiv.org/abs/2609.29587

[h] Back to Home