NeFut Logo NeFut
Admin Login

[CS.AI] Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models

Published at: 2026-09-11 22:00 Last updated: 2026-09-12 06:35
#algorithm #AI #Machine Learning

Speech-to-speech (S2S) models are now deployed in dubbing, translation and voice agents. Unlike pure‑text models, they receive the speaker's voice, which carries gender cues. A faithful system should identify gender from acoustic cues rather than from stereotypical content.

Most existing S2S systems generate output in a single, fixed voice, making it hard to detect whether content bias drifts the perceived gender of the rendered voice. Even if the voice stays constant, bias may still appear in how the model attributes gender.

This work asks two questions: 1) When the model re‑speaks the input, does gender‑stereotyped wording shift the perceived gender of the output voice (voice rendering)? 2) When the model explicitly states the speaker's gender, does it follow the voice or the content (gender attribution)?

A controlled experiment crosses male and female source voices with masculine, neutral and feminine passages, covering five open‑ and closed‑source models in English, Spanish and Mandarin. The rendered voice shows virtually no drift – the acoustic identity is preserved. However, all models decide gender from the textual content. Moving the passage from masculine → neutral → feminine multiplies the odds of a "female" judgment by 1.7‑24×.

When content conflicts with voice, the worst model misgenders 90% of the time; when they align, the error rate drops to 2%. Thus bias hides in gender attribution, invisible to fixed‑voice evaluations, and must be audited as S2S systems increasingly speak for real people.

Review

Original Source: https://arxiv.org/abs/2609.09263

[h] Back to Home