Search arXivSearch

arXiv · 2401.10242

DanceMeld: Unraveling Dance Phrases with Hierarchical Latent Codes for Music-to-Dance Synthesis

Abstract

In the realm of 3D digital human applications, music-to-dance presents a challenging task. Given the one-to-many relationship between music and dance, previous methods have been limited in their approach, relying solely on matching and generating corresponding dance movements based on music rhythm. In the professional field of choreography, a dance phrase consists of several dance poses and dance movements. Dance poses composed of a series of basic meaningful body postures, while dance movements can reflect dynamic changes such as the rhythm, melody, and style of dance. Taking inspiration from these concepts, we introduce an innovative dance generation pipeline called DanceMeld, which comprising two stages, i.e., the dance decouple stage and the dance generation stage. In the decouple stage, a hierarchical VQ-VAE is used to disentangle dance poses and dance movements in different feature space levels, where the bottom code represents dance poses, and the top code represents dance movements. In the generation stage, we utilize a diffusion model as a prior to model the distribution and generate latent codes conditioned on music features. We have experimentally demonstrated the representational capabilities of top code and bottom code, enabling the explicit decoupling expression of dance poses and dance movements. This disentanglement not only provides control over motion details, styles, and rhythm but also facilitates applications such as dance style transfer and dance unit editing. Our approach has undergone qualitative and quantitative experiments on the AIST++ dataset, demonstrating its superiority over other methods.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xin Gao, Li Hu, Peng Zhang, Bang Zhang, Liefeng Bo. 2023-11-30. DanceMeld: Unraveling Dance Phrases with Hierarchical Latent Codes for Music-to-Dance Synthesis. https://arxiv.org/abs/2401.10242

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Bidirectional Temporal Dynamics Modeling for EEG-based Driving Fatigue Recognition

Driving fatigue is a major contributor to traffic accidents and poses a serious threat to road safety. Electroencephalography (EEG) provides a direct measurement of neural activity, yet EEG-based fatigue recognition is hindered by strong non-stationarity and asymmetric neural dynamics. To address these challenges, we propose DeltaGateNet, a novel framework that explicitly captures Bidirectional temporal dynamics for EEG-based driving fatigue recognition. Our key idea is to introduce a Bidirectional Delta module that decomposes first-order temporal differences into positive and negative components, enabling explicit modeling of asymmetric neural activation and suppression patterns. Furthermore, we design a Gated Temporal Convolution module to capture long-term temporal dependencies for each EEG channel using depthwise temporal convolutions and residual learning, preserving channel-wise specificity while enhancing temporal representation robustness. Extensive experiments conducted under both intra-subject and inter-subject evaluation settings on the public SEED-VIG and SADT driving fatigue datasets demonstrate that DeltaGateNet consistently outperforms existing methods. On SEED-VIG, DeltaGateNet achieves an intra-subject accuracy of 81.89% and an inter-subject accuracy of 55.55%. On the balanced SADT 2022 dataset, it attains intra-subject and inter-subject accuracies of 96.81% and 83.21%, respectively, while on the unbalanced SADT 2952 dataset, it achieves 96.84% intra-subject and 84.49% inter-subject accuracy. These results indicate that explicitly modeling Bidirectional temporal dynamics yields robust and generalizable performance under varying subject and class-distribution conditions.

cs.OH

Fixing ill-formed UTF-16 strings with SIMD instructions

UTF-16 is a widely used Unicode encoding representing characters with one or two 16-bit code units. The format relies on surrogate pairs to encode characters beyond the Basic Multilingual Plane, requiring a high surrogate followed by a low surrogate. Ill-formed UTF-16 strings -- where surrogates are mismatched -- can arise from data corruption or improper encoding, posing security and reliability risks. Consequently, programming languages such as JavaScript include functions to fix ill-formed UTF-16 strings by replacing mismatched surrogates with the Unicode replacement character (U+FFFD). We propose using Single Instruction, Multiple Data (SIMD) instructions to handle multiple code units in parallel, enabling faster and more efficient execution. Our software is part of the Google JavaScript engine (V8) and thus part of several major Web browsers.

cs.OH

MRSeqStudio: MRI Sequence Design and Simulation as a Service in a Free and Open-Source Web Platform

MRI sequence prototyping increasingly relies on graphical design environments and numerical simulators to accelerate development and validation. While several platforms support interactive sequence construction, fully web-based solutions that combine integrated phantom management, high-fidelity Bloch simulation, and scalable multi-user deployment remain limited. We present MRSeqStudio, a web-based platform for interactive MR sequence design and simulation. The tool adopts a block-based representation model with real-time visualization and native JSON/Pulseq export. Simulations are performed using the GPU-enabled Bloch simulator KomaMRI, which enables accurate modeling of arbitrary pulse sequences and phantoms within an installation-free architecture. The system separates front-end interaction from back-end simulation services to support concurrent multi-user access. Sequence validity was assessed by comparing GRE and bSSFP implementations against equivalent sequences designed in mtrk and gammaSTAR. The resulting images showed minimal absolute differences and high mean structural similarity indices (SSIM). Stress testing under burst-request conditions demonstrated stable performance with up to 100 concurrent users on a high-performance desktop deployment. A comparative workflow analysis with mtrk and gammaSTAR further examined differences in representation models, parameter propagation strategies, and integration levels across platforms, highlighting the relative strengths and limitations of each tool. Results indicate that MRSeqStudio provides a reliable and accessible environment for MR sequence prototyping, combining web-native deployment with Bloch-level simulation fidelity and integrated phantom visualization.

cs.OH