arXiv · 2609.37400
BeatDance: Generating Beat-Consistent 3D Dance with Hierarchical Spatial-Temporal Modeling
Abstract
Generating realistic 3D dance from music is a challenging task that requires accurate synchronization with musical rhythms while capturing the spatial complexity of human motion. Although existing methods can generate physically plausible dance motions, they often struggle to achieve precise alignment with music, such as the beat. To address this limitation, we propose a novel diffusion-based framework, BeatDance, with two components: 1) We present a Hierarchical Decoupled Attention (HDA) module, which first disentangles the learning of human pose and temporal dynamics. A hierarchical structure is then employed to capture both short-term and long-term dependencies, thereby enhancing spatial-temporal modeling. 2) We adopt cycle-consistent learning by introducing an auxiliary dance-to-music module. During training, discrepancies between the reconstructed and original music induce a stronger loss signal, effectively encouraging the consistency property between the music and dance motion. Extensive experimental results demonstrate that our proposed approach outperforms recent competitive methods on two benchmark datasets.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xiaojian Shen, Dahu Shi, Jianrong Zhang, Hai Li, Hongwei Zhao, Dawei Zhang, Yunzhi Zhuge, Zhiliang Wu, Guanghui Yue, Wei Zhou. 2026-09-29. BeatDance: Generating Beat-Consistent 3D Dance with Hierarchical Spatial-Temporal Modeling. https://doi.org/10.1016/j.patcog.2026.114344
Cite the original work for its findings. Save a collection to share your selection of sources.