NeFut Logo NeFut
中 Admin Login

[CS.AI] How Much Can Reliability Drift Under a Fixed Confidence Distribution?

Published at: 2026-10-01 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #optimization

A classifier’s conditional accuracy may change while its confidence distribution stays exactly the same. We investigate the worst‑case movement of the reliability relation under covariate shifts that preserve the confidence‑score distribution, constraining the re‑weighting within each confidence level by a $\chi^2$ budget.

On an interval of budgets that can be computed from the source distribution, the worst case equals the square root of the budget times the within‑level variance of the correctness propensity – the grouping‑loss term in calibration‑refinement decompositions. Beyond this interval the profile is governed by the tails of the propensity law, and the entire upward profile determines the centred within‑level law; consequently, calibration residual and grouping variance do not generally determine fragility, except when labels and predictions are deterministic.

Because the propensity is unobserved, we restrict re‑weightings to a learned finite readout within confidence bins, bound the missed part by the grouping variance that remains inside readout cells, estimate the restricted profile with role‑separated labels, and provide a separate split‑sample lower confidence bound.

Experiments on ImageNet show that for four of six primary classifiers and nine of twelve additional released models, the bound is positive in both splits; after temperature scaling it remains positive for three of eighteen models. Optimised re‑weightings fitted without evaluation labels track the estimated profile on held‑out drift; a label‑permutation diagnostic yields near‑zero agreement for this statistic while largely reproducing the correlation observed for unsigned random re‑weightings.

The conclusion is that reliability drift fragility can be described by the product of the budget’s square root and the within‑level variance, and it can be reliably estimated in practice using finite readouts and split‑sample bounds.

Review

Original Source: https://arxiv.org/abs/2609.38917

[h] Back to Home