NeFut Logo NeFut
Admin Login

[CS.AI] Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation

Published at: 2026-09-21 22:00 Last updated: 2026-09-22 02:29
#AI #Machine Learning #Neural

We introduce ReACT‑TTS, a two‑stage framework for conversational speech generation. In the first stage, a one‑second pre‑response facial sequence of the listener is used to predict the emotion and prosody of the upcoming utterance; the second stage injects these predictions into a Grad‑TTS backbone for end‑to‑end speech synthesis.

Under a strict dyadic MELD protocol, temporal conditioning improves mean macro‑F1 and VAD concordance across ten random seeds compared to a text‑only baseline, while overall accuracy remains essentially unchanged. Ablation studies reveal that temporal modeling outperforms all visual variants, that an explicit early‑to‑late distinction is unnecessary, and that correct listener reactions beat cyclic mismatches on average.

In a contextual‑appropriateness study with 20 speech researchers, 76% preferred the temporal model, 9% favored the text‑only version, and 15% expressed no clear preference.

The source code is released at https://github.com/CYJ1/ReACT-TTS_public.

Review: The findings demonstrate that immediate listener facial dynamics serve as complementary cues for response planning, offering a novel visual signal for emotion‑aware TTS.

Original Source: https://arxiv.org/abs/2609.21683

[h] Back to Home