NeFut Logo NeFut
中 Admin Login

[CS.AI] Beyond Answer Confidence: A Controlled Audit of Self‑Knowledge in a Black‑Box Decision Model

Published at: 2026-10-03 22:00 Last updated: 2026-10-06 12:11
#Machine Learning #LLM #Artificial Intelligence

Decision models emit probabilities for routing, abstention, and automated actions. Calibration makes these probabilities useful on average, yet it does not reveal whether low confidence stems from chance or missing knowledge, nor whether confidence drops when the model ventures beyond its known domain. We performed a controlled audit on Jev, a decision model, covering over 15 public datasets and six generated task families, with paired interventions that vary the amount of information supplied for a fixed item.

Jev’s confidence is well‑calibrated on familiar closed‑choice tasks, but when answer‑relevant information is absent it still assigns up to 0.80 to a salient option. On news items beyond an observed knowledge boundary, confidence exceeds actual accuracy by 0.21–0.33, and recalibrating on earlier months does not close this gap.

Targeted yes/no questions provide sharper diagnostics: determining whether an outcome is settled yields an AUROC of 1.00, and assessing whether evidence suffices yields an AUROC of 0.95, compared with 0.85 for raw confidence on the same items. Asking Jev if it knows the answer flags fabricated entities and post‑boundary news (AUROC 0.91), but with realistic names or dates removed the advantage disappears, performing no better than ordinary answer uncertainty.

Thus, black‑box knowledge audits require explicit controls for surface cues to avoid misleading conclusions. The implementation is available at https://github.com/Syntheme/beyond-answer-confidence.

Review

Original Source: https://arxiv.org/abs/2610.01006

[h] Back to Home