NeFut Logo NeFut
Admin Login

[CS.AI] Talking Head Synthesis with Facial Landmark Guidance via 3D Gaussian Splatting

Published at: 2026-09-17 22:00 Last updated: 2026-09-18 00:46
#AI #Machine Learning #Neural

Audio-driven digital human generation is essential for virtual communication, immersive interaction, and media production. With the rise of Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), recent talking‑head systems achieve more faithful 3D facial geometry and appearance. A core challenge is that speech features capture temporal acoustic patterns but lack explicit facial layout cues, so driving 3D facial deformation directly from audio can cause inaccurate mouth motion, weak expression details, and local artifacts. To tackle this, we introduce a facial‑keypoint‑guided spatial enhancement module. Predicted landmarks supply structural cues for selecting and enriching spatial points around expression‑sensitive regions. We also add a global landmark compensation mechanism that encodes the full set of keypoints into a conditioning vector to refine 3DGS attributes, providing whole‑face structural information to the underlying shape representation. Experiments under self‑driven and cross‑driven settings demonstrate improvements in visual quality, facial realism, and lip synchronization.

Review

Original Source: https://arxiv.org/abs/2609.17422

[h] Back to Home