NeFut Logo NeFut
Admin Login

[CS.AI] The Privacy‑Hallucination Tradeoff in Differentially Private Language Models

Published at: 2026-09-02 22:00 Last updated: 2026-09-03 02:56
#AI #Machine Learning #Artificial Intelligence

In high‑stakes domains such as healthcare, protecting privacy and ensuring factual accuracy are equally critical. We uncover and systematically investigate a privacy‑hallucination tradeoff in differentially private (DP) language models. Empirical results show that models pre‑trained or fine‑tuned with DP generate more hallucinations than non‑DP counterparts, and the frequency and severity of hallucinations increase as the privacy budget (ε) becomes stricter. Further analysis reveals that DP mechanisms flatten output distributions, which can shift probability mass toward factually incorrect alternatives. To examine the role of information frequency, we control the occurrence of facts in the training data; higher fact frequency mitigates hallucination risk in DP models. Overall, our findings highlight the need for more nuanced privacy‑preserving techniques that provide rigorous guarantees without sacrificing factual correctness.

Review

Original Source: https://arxiv.org/abs/2609.00492

[h] Back to Home