NeFut Logo NeFut
Admin Login

[CS.AI] PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads

Published at: 2026-08-07 22:00 Last updated: 2026-08-08 01:08
#AI #Machine Learning #LLM

A recent study proposes a novel approach called Phoneme-Driven Gaussian Splatting (PD-GS) for more realistic lip motion. Traditional 3D Gaussian Splatting (3DGS) methods can quickly render photorealistic talking heads but still struggle with lip articulation, resulting in over-smoothed mouth movements that may violate hard articulatory constraints. To address this, the researchers introduced the PD-GS method, which leverages an automatic speech recognition (ASR) and forced-alignment pipeline to obtain time-aligned phoneme tokens and combine them with a 3DGS talker. The core component of PD-GS is the Linguistic Fusion Module (LFM), which adaptively fuses continuous audio context with discrete phoneme embeddings through a learned gate, preserving smooth audio-driven dynamics while strengthening phoneme guidance on articulation-critical segments. The experiments demonstrate that PD-GS achieves the best lip geometry among the compared baselines on the HDTF dataset and qualitatively reduces closure violations in challenging phoneme sequences, yielding more linguistically faithful neural avatars. Blogger's Review: This paper presents an innovative PD-GS algorithm that combines phoneme tokens with a 3DGS model to achieve more realistic lip motion and lip geometry. This approach has the potential to be applied in virtual avatars, audio-driven animation, and other related fields.

Original Source: https://arxiv.org/abs/2608.05218

[h] Back to Home