Search arXiv⌕ Search

arXiv · 2610.09846

DynaConTalk: Wavelet-Constrained Diffusion for Long-Form and Controllable Holistic Co-Speech 3D Motion

Abstract

Holistic co-speech animation is prone to averaging in both motion representation and speech conditioning. In coordinate-space diffusion, slow body posture, mid-frequency gesture strokes, and fast hand or facial details are entangled in one prediction target, often producing low-variance, over-smoothed motion. Meanwhile, dense rhythmic and acoustic cues can dominate sparse content-specific information under fixed multimodal fusion. We present DynaConTalk, a wavelet-constrained diffusion framework for long-form and controllable holistic co-speech motion generation. Diffusion operates in stationary wavelet transform (SWT) coefficient space, whose temporally aligned bands separate coarse posture evolution, gesture strokes, and fine expressive details. Our dynamic gating network preserves HuBERT and speaker identity as a base and selectively adds rhythm, mel, and transcript features through motion-state- and noise-aware residual gates. Attention pooling and learned depth routing deliver complementary conditions to each denoising stage, while a frame-resolution rhythm path preserves precise timing. A signed proposal-consensus update then reconciles these conditions with the evolving motion state. Matched-noise constraint injection uses the same sampling interface for history continuation and localized keypose repair, and extends to reference-guided control. Separate body-hand and facial denoisers, followed by inverse SWT and a pose-driven root regressor, produce holistic motion. Experiments evaluate generation quality, facial accuracy, temporal continuity, and controllable editing. Code, models, and the interactive editing interface are available at https://github.com/zhuyifeiabcd1/DynaConTalk.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yifei Zhu, Yangyang Cai, Mingyi Shi, Miao Cheng, Lin Gu, Taku Komura, Yoshifumi Kitamura. 2026-10-07. DynaConTalk: Wavelet-Constrained Diffusion for Long-Form and Controllable Holistic Co-Speech 3D Motion. https://arxiv.org/abs/2610.09846

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

HistCAD: Constraint-Aware Parametric CAD Histories for Evaluating Editability

Sketch constraints specify geometric conditions for constructing and modifying parametric CAD models. We study whether predicted constraints allow a given history to reproduce the required initial model and support prescribed dimensional edits. We introduce HistCAD, an executable representation and dataset whose Academic and Industrial collections contain 180,495 parametric construction histories with entity-referenced sketch constraints and retained feature operations. Predictors receive these histories with the geometry and feature definitions retained and explicit sketch constraints removed. They generate constraints for every sketch without seeing the edit request. The benchmark compares models built with alternative constraint sets for the same history under the same dimensional edit. An edit succeeds when the model reproduces the required initial geometry, reaches the target value, preserves specified relations and unedited dimensions in the target sketch, and rebuilds through the complete history. A predictor trained on both collections and supplied with descriptions of the input histories achieves overall edit success of 52.4% on Academic and 29.0% on Industrial. Models retaining only endpoint-connectivity constraints in the target sketch and the history's constraints elsewhere can reach the target and rebuild while failing preservation. For all-sketch predictions, we retain the target-sketch prediction and restore the history's constraints in other sketches. More models then reproduce the required initial geometry, and some of these newly matched models complete the edit. HistCAD connects constraint learning to the construction and revision of parametric CAD models.

cs.GR↗

AnchorFlow: Learning Anchor Placement for Faithful and Editable SVG Reconstruction

Raster-to-SVG reconstruction requires faithful geometry and a compact control structure for editing. A central challenge is deciding where to place anchors: raster appearance alone does not determine how a contour should be divided into Bézier segments. We present AnchorFlow, which learns anchor placement from designer-authored SVGs to reconstruct accurate curves with sparse controls. Our key idea is a sparse anchor field that jointly encodes contour geometry and reference segment junctions, including those along smooth contours. An anchor decoder predicts explicit anchor proposals from features learned under field supervision. These proposals guide boundary-constrained fitting and local refinement to recover cubic Bézier paths. On clean single-path inputs, AnchorFlow achieves 99.52% mean IoU while using 56.6% fewer anchors on average than AdaVec, with lower boundary error and closer agreement with source-SVG anchor layouts. Under boundary perturbations, it maintains high fidelity with limited anchor growth. Integrated into a component-based pipeline, the same path module also produces compact, faithful full-image reconstructions. On four local-editing tasks, our outputs require less median active time and fewer actions than AdaVec and LIVE while retaining high target-shape accuracy.

cs.GR↗

Tabula Rasa: Monte Carlo estimation of unit-variance noise with controlled spatio-temporal correlation

We suggest a method to generate time-varying Gaussian noise with controlled variance and controlled temporal correlation. This noise is used in several downstream tasks for temporal control and temporal coherence. The core technical idea is to phrase this problem as joint Monte-Carlo estimation of both a classic pixel reconstruction and estimation of variance using the concept of "sketching" from the database literature. We demonstrate that our method allows temporal control for downstream tasks with simpler and faster code than previous methods.

cs.GR↗