NeFut Logo NeFut
中 Admin Login

[CS.AI] Which Influence Are We Estimating? The Role of Counterfactual Specifications in Data Attribution

Published at: 2026-09-29 22:00 Last updated: 2026-09-30 01:41
#AI #Machine Learning #LLM

Estimating how individual training examples influence model behavior underlies data debugging, valuation, and attribution. Existing influence estimators often produce conflicting rankings, typically blamed on approximation error. We argue that a deeper source of disagreement is specification mismatch: influence depends on (1) the behavior being attributed, (2) the intervention applied to each training example, and (3) the counterfactual training process that maps the intervention to a model response. These choices become critical when the target behavior requires a tractable surrogate such as query loss, a logit, or a margin. We formalize influence as a counterfactual estimand, separate specification mismatch across estimands from approximation error in estimating a fixed estimand, and organize representative estimators by their implied specifications. We also derive a local decomposition that reveals how behavior signals, training signals, and counterfactual parameter responses interact. Controlled experiments show that exact estimands under different specifications can yield different rankings, while approximation error grows as perturbations move farther from their linearization points. Experiments on noisy‑label detection and LLM attribution demonstrate that specification choices markedly affect attribution quality, especially the choice of behavior surrogate. Behavior‑aligned specifications can uncover target‑specific training examples hidden by default loss‑based or similarity‑based specifications. These findings establish specification analysis as a necessary first step for interpreting and comparing data influence estimators.

Review

Original Source: https://arxiv.org/abs/2609.31214

[h] Back to Home