NeFut Logo NeFut
Admin Login

[CS.AI] Anonymization, Not Elimination: Utility-Preserved Speech Anonymization

Published at: 2026-09-05 22:00 Last updated: 2026-09-06 01:02
#Machine Learning #Neural #Artificial Intelligence

The rapid growth of large‑scale speech corpora has made privacy protection a pressing issue. Existing anonymization techniques often disrupt acoustic continuity or reduce vocal diversity, which harms downstream tasks such as automatic speech recognition, text‑to‑speech synthesis, and speech emotion recognition. To preserve both privacy and utility, we introduce a two‑stage framework. In the first stage, a generative speech editing model seamlessly replaces personally identifiable information (PII), ensuring content privacy. In the second stage, we propose F3‑VA, a flow‑matching based anonymization system with a three‑stage design that generates diverse and distinguishable anonymized speakers, protecting voice identity. For evaluation, we complement traditional acoustic‑ and content‑based speaker verification metrics with utility assessment by training ASR, TTS, and SER models from scratch. Experiments show that our approach outperforms VoicePrivacy Challenge baselines in privacy protection while incurring minimal utility loss, and the new protocol offers a more realistic view of anonymized speech utility.

Review: This work combines generative editing and flow‑matching to achieve simultaneous content and voice anonymization, and validates utility through end‑to‑end training of downstream models, presenting a robust and practical solution for speech privacy.

Original Source: https://arxiv.org/abs/2604.17000

[h] Back to Home