Search arXivSearch

arXiv · 2504.13386

Supervising 3D Talking Head Avatars with Analysis-by-Audio-Synthesis

Abstract

In order to be widely applicable, speech-driven 3D head avatars must articulate their lips in accordance with speech, while also conveying the appropriate emotions with dynamically changing facial expressions. The key problem is that deterministic models produce high-quality lip-sync but without rich expressions, whereas stochastic models generate diverse expressions but with lower lip-sync quality. To get the best of both, we seek a stochastic model with accurate lip-sync. To that end, we develop a new approach based on the following observation: if a method generates realistic 3D lip motions, it should be possible to infer the spoken audio from the lip motion. The inferred speech should match the original input audio, and erroneous predictions create a novel supervision signal for training 3D talking head avatars with accurate lip-sync. To demonstrate this effect, we propose THUNDER (Talking Heads Under Neural Differentiable Elocution Reconstruction), a 3D talking head avatar framework that introduces a novel supervision mechanism via differentiable sound production. First, we train a novel mesh-to-speech model that regresses audio from facial animation. Then, we incorporate this model into a diffusion-based talking avatar framework. During training, the mesh-to-speech model takes the generated animation and produces a sound that is compared to the input speech, creating a differentiable analysis-by-audio-synthesis supervision loop. Our extensive qualitative and quantitative experiments demonstrate that THUNDER significantly improves the quality of the lip-sync of talking head avatars while still allowing for generation of diverse, high-quality, expressive facial animations. The code and models will be available at https://thunder.is.tue.mpg.de/

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Radek Daněček, Carolin Schmitt, Senya Polikovsky, Michael J. Black. 2026-01-27. Supervising 3D Talking Head Avatars with Analysis-by-Audio-Synthesis. https://arxiv.org/abs/2504.13386

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Opacity Is Not Just Opacity

Web graphics travel with content across pages and themes, where changing backgrounds can require recoloring and maintenance. Opacity already makes a fixed object's appearance depend on its background, yet is usually understood only as how much the object obscures it. In fact, opacity controls the scaling of the object-background color difference; transparency is only one effect of this relationship. Zero places the output at the background and one at the source color, but difference scaling need not stop at either position. We retain the compositing expression and extend the coefficient domain from $[0,1]$ to the real numbers: negative values reverse the difference, whereas values above one expand it in the same direction. We focus on same-direction expansion for reusing Web graphics across backgrounds. Each object carries a fixed source color and coefficient, while the actual background determines the enhancement direction. Background-adaptive contrast enhancement thus becomes part of the object's compositing properties, reducing the design and maintenance of separate color variants. The implementation reuses the original equation without increasing the per-pixel arithmetic operation count within the same pipeline. Enumerating all 8-bit sRGB source colors on 16 predefined light and dark canvases, a fixed $α=1.1$ increases the contrast ratio in 99.8145% of combinations. Without changing source colors, 4.8346% of all combinations newly reach the $3:1$ contrast threshold. Output validation and timing across three browsers demonstrate implementation in the same WebGL pipeline, with no sustained additional runtime observed.

cs.GR

Constrained Program Generation for 3D Reaction Animation with a 0.8B Model

Visualizing a chemical reaction requires making its molecular changes visible while keeping the animation faithful to the stated chemistry. Equations, structural diagrams and molecular viewers provide complementary descriptions, but assembling an interactive three-dimensional explanation still requires specifying the changes and checking their consistency. We present ChemXRG, a domain-specific language (DSL) framework that addresses this gap by representing a reaction animation as an executable program. Persistent atom identifiers and explicit bond, charge and grouping operations connect the symbolic reaction to the displayed transformation. This shared representation lets generation, verification and rendering operate on the same account of what changes. Given known, atom-mapped reactant and product structures, reaction-grounded constraints fix input-determined facts and restrict action choices; execution checks validate the resulting transformation before geometry and frames are constructed. We implement this paradigm with a reaction-program corpus and ChemQwen, a trained 0.8B DSL generator. Paired and component evaluations show improved compiler acceptance and normalized full-program agreement under input-conditioned constraints, while identifying remaining failures that require execution checks. A public browser application demonstrates the connection from symbolic reaction descriptions to inspectable programs and interactive 3D animations.

cs.GR

NaRPA: Navigation and Rendering Pipeline for Astronautics

This paper presents the applications of scientific ray-tracing in modeling and simulating light transport for space-borne image data generation. A ray-tracing engine, the Navigation and Rendering Pipeline for Astronautics (NaRPA), is introduced as a rendering framework to generate virtual datasets and support simulations for robust navigation pipelines. Sensor and environment models that enable the synthesis of space-to-space and ground-to-space virtual observations are presented. The work demonstrates the capabilities of simulating passive and active vision-based sensors using NaRPA to facilitate the design, testing, and verification of aerospace visual navigation algorithms. Additionally, the paper describes a velocimeter LiDAR model and its statistical validation with experimental data.

cs.GR