Search arXivSearch

arXiv subjects

Rui Hong

Publications and source records attributed to Rui Hong.

10 recordsLinked to original sources

Conditional Collapse in Sign Language Production: A Diagnostic and a Scaling Argument

Sign Language Production (SLP) is the task of generating avatar sign language motion from natural language text. The quality of the generated motion is typically evaluated by a motion-space Fr\'echet distance (FID) and back-translation (BT) BLEU score on benchmarks such as How2Sign. Both metrics can improve substantially while the underlying generator fails to faithfully represent the sign language gestures. In this work we propose to evaluate the generated motion at three independent levels: ($\tau1$) initial-pose conditioning, ($\tau2$) output diversity, and ($\tau3$) target faithfulness. We compute these as pairwise-distance ratios using latent representations of a frozen motion autoencoder (MoAE). We evaluate 14 SLP model checkpoints on the How2Sign dataset, including a re-implemented Neural Sign Actors (NSA), and show that $\tau3$ faithfulness is never attained, while FID varies by nearly two orders of magnitude and is uncorrelated with faithfulness. We show that on the isolated gloss dataset ASL3DWord favorable $\tau3$ can be attained, hence isolating the size of the sentence-level paired-dataset as the bottleneck.

cs.CV

To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs

When VLMs answer correctly, do they genuinely rely on visual information? We introduce a Tri-Layer Diagnostic Framework with three per-sample metrics: Latent Anomaly Detection, Visual Necessity Score, and Competition Score, which disentangle perception, dependency, and alignment failures. Across 9 VLMs and 9,000 model-sample pairs under counterfactual blind, noise, and conflict interventions, 72.9% of samples exhibit Visual Sycophancy, a Split Beliefs pattern in which internal evidence is preserved yet a hallucinated answer is decoded, while zero samples show Robust Refusal, indicating that current alignment training has eliminated refusal as a decoding outcome. Scaling within the Qwen-VL family, both within- and across-generation, monotonically reduces Language Shortcuts but amplifies Visual Sycophancy, showing that scale and newer post-training alone cannot resolve the grounding problem. Diagnostic scores further enable a training-free selective-prediction strategy yielding up to +9.5 percentage points accuracy at 50% coverage.

cs.CV

Toward Phonology-Guided Sign Language Motion Generation: A Diffusion Baseline and Conditioning Analysis

Generating natural, correct, and visually smooth 3D avatar sign language motion conditioned on the text inputs continues to be very challenging. In this work, we train a generative model of 3D body motion and explore the role of phonological attribute conditioning for sign language motion generation, using ASL-LEX 2.0 annotations such as hand shape, hand location and movement. We first establish a strong diffusion baseline using an Human Motion MDM-style diffusion model with SMPL-X representation, which outperforms SignAvatar, a state-of-the-art CVAE method, on gloss discriminability metrics. We then systematically study the role of text conditioning using different text encoders (CLIP vs. T5), conditioning modes (gloss-only vs. gloss+phonological attributes), and attribute notation format (symbolic vs. natural language). Our analysis reveals that translating symbolic ASL-LEX notations to natural language is a necessary condition for effective CLIP-based attribute conditioning, while T5 is largely unaffected by this translation. Furthermore, our best-performing variant (CLIP with mapped attributes) outperforms SignAvatar across all metrics. These findings highlight input representation as a critical factor for text-encoder-based attribute conditioning, and motivate structured conditioning approaches where gloss and phonological attributes are encoded through independent pathways.

cs.CV

Gesture-Aware Pretraining and Token Fusion for 3D Hand Pose Estimation

Estimating 3D hand pose from monocular RGB images is fundamental for applications in AR/VR, human-computer interaction, and sign language understanding. In this work we focus on a scenario where a discrete set of gesture labels is available and show that gesture semantics can serve as a powerful inductive bias for 3D pose estimation. We present a two-stage framework: gesture-aware pretraining that learns an informative embedding space using coarse and fine gesture labels from InterHand2.6M, followed by a per-joint token Transformer guided by gesture embeddings as intermediate representations for final regression of MANO hand parameters. Training is driven by a layered objective over parameters, joints, and structural constraints. Experiments on InterHand2.6M demonstrate that gesture-aware pretraining consistently improves single-hand accuracy over the state-of-the-art EANet baseline, and that the benefit transfers across architectures without any modification.

cs.CV

Motion-Adaptive Temporal Attention for Lightweight Video Generation with Stable Diffusion

We present a motion-adaptive temporal attention mechanism for parameter-efficient video generation built upon frozen Stable Diffusion models. Rather than treating all video content uniformly, our method dynamically adjusts temporal attention receptive fields based on estimated motion content: high-motion sequences attend locally across frames to preserve rapidly changing details, while low-motion sequences attend globally to enforce scene consistency. We inject lightweight temporal attention modules into all UNet transformer blocks via a cascaded strategy -- global attention in down-sampling and middle blocks for semantic stabilization, motion-adaptive attention in up-sampling blocks for fine-grained refinement. Combined with temporally correlated noise initialization and motion-aware gating, the system adds only 25.8M trainable parameters (2.9\% of the base UNet) while achieving competitive results on WebVid validation when trained on 100K videos. We demonstrate that the standard denoising objective alone provides sufficient implicit temporal regularization, outperforming approaches that add explicit temporal consistency losses. Our ablation studies reveal a clear trade-off between noise correlation and motion amplitude, providing a practical inference-time control for diverse generation behaviors.

cs.CV

Hybrid-Precision Block-Jacobi Preconditioned GMRES Solver for Linear System in Circuit Simulation

As integrated circuits become increasingly complex, the demand for efficient and accurate simulation solvers continues to rise. Traditional solvers often struggle with large-scale sparse systems, leading to prolonged simulation times and reduced accuracy. In this paper, a hybrid-precision block-Jacobi preconditioned GMRES solver is proposed to solve the large sparse system in circuit simulation. The proposed method capitalizes on the structural sparsity and block properties of circuit matrices, employing a novel hybrid-precision strategy that applies single-precision arithmetic for computationally intensive tasks and double-precision arithmetic for critical accuracy-sensitive computations. Additionally, we use the graph partitioning tools to assist in generating preconditioners, ensuring an optimized preconditioning process. For large-scale problems, we adopt the restart strategy to increase the computational efficiency. Through rigorous mathematical reasoning, the convergence and error analysis of the proposed method are carried out. Numerical experiments on various benchmark matrices demonstrate that our approach significantly outperforms existing solvers, including SuperLU, KLU, and SFLU, in terms of both preconditioning and GMRES runtime. The proposed hybrid-precision preconditioner effectively improves spectral clustering, leading to faster solutions.

math.NA

Entanglement scaling and criticality of infinite-size quantum many-body systems in continuous space addressed by a tensor network approach

Simulating strongly-correlated quantum systems in continuous space belongs to the most challenging and long-concerned issues in quantum physics. This work investigates the quantum entanglement and criticality of the ground-state wave-functions of infinitely-many coupled quantum oscillators (iCQOs). The essential task involves solving a set of partial differential equations (Schr\"odinger equations in the canonical quantization picture) with infinitely-many variables, which currently lacks valid methods. By extending the imaginary-time evolution algorithm with translationally-invariant functional tensor network, we simulate the ground state of iCQOs with the presence of two- and three-body couplings. We determine the range of coupling strengths where there exists a real ground-state energy (dubbed as physical region). With two-body couplings, we reveal the logarithmic scaling law of entanglement entropy (EE) and the polynomial scaling law of correlation length against the virtual bond dimension $\chi$ at the dividing point of physical and non-physical regions. These two scaling behaviors are signatures of criticality, according to the previous results in quantum lattice models, but were not reported in continuous-space quantum systems. The scaling coefficients result in a central charge $c=1$, indicating the presence of free boson conformal field theory (CFT). We further show that the presence of three-body couplings, for which there are no analytical or numerical results, breaks down the CFT description at the dividing point. Our work reveals the scaling behaviors of EE in continuous-space quantum many-body systems. These results provide strong numerical evidence supporting the efficiency of TN in representing continuous-space quantum wave-functions in the thermodynamic limit and offer an efficient approach to studying entanglement properties and criticality in continuous space.

quant-ph

Functional Tensor Network Solving Many-body Schr\"odinger Equation

Schr\"odinger equation belongs to the most fundamental differential equations in quantum physics. However, the exact solutions are extremely rare, and many analytical methods are applicable only to the cases with small perturbations or weak correlations. Solving the many-body Schr\"odinger equation in the continuous spaces with the presence of strong correlations is an extremely important and challenging issue. In this work, we propose the functional tensor network (FTN) approach to solve the many-body Schr\"odinger equation. Provided the orthonormal functional bases, we represent the coefficients of the many-body wave-function as tensor network. The observables, such as energy, can be calculated simply by tensor contractions. Simulating the ground state becomes solving a minimization problem defined by the tensor network. An efficient gradient-decent algorithm based on the automatically differentiable tensors is proposed. We here take matrix product state (MPS) as an example, whose complexity scales only linearly with the system size. We apply our approach to solve the ground state of coupled harmonic oscillators, and achieve high accuracy by comparing with the exact solutions. Reliable results are also given with the presence of three-body interactions, where the system cannot be decoupled to isolated oscillators. Our approach is simple and with well-controlled error, superior to the highly-nonlinear neural-network solvers. Our work extends the applications of tensor network from quantum lattice models to the systems in the continuous space. FTN can be used as a general solver of the differential equations with many variables. The MPS exemplified here can be generalized to, e.g., the fermionic tensor networks, to solve the electronic Schr\"odinger equation.

quant-ph

Predicting Quantum Potentials by Deep Neural Network and Metropolis Sampling

The hybridizations of machine learning and quantum physics have caused essential impacts to the methodology in both fields. Inspired by quantum potential neural network, we here propose to solve the potential in the Schrodinger equation provided the eigenstate, by combining Metropolis sampling with deep neural network, which we dub as Metropolis potential neural network (MPNN). A loss function is proposed to explicitly involve the energy in the optimization for its accurate evaluation. Benchmarking on the harmonic oscillator and hydrogen atom, MPNN shows excellent accuracy and stability on predicting not just the potential to satisfy the Schrodinger equation, but also the eigen-energy. Our proposal could be potentially applied to the ab-initio simulations, and to inversely solving other partial differential equations in physics and beyond.

quant-ph

Automatically Differentiable Quantum Circuit for Many-qubit State Preparation

Constructing quantum circuits for efficient state preparation belongs to the central topics in the field of quantum information and computation. As the number of qubits grows fast, methods to derive large-scale quantum circuits are strongly desired. In this work, we propose the automatically differentiable quantum circuit (ADQC) approach to efficiently prepare arbitrary quantum many-qubit states. A key ingredient is to introduce the latent gates whose decompositions give the unitary gates that form the quantum circuit. The circuit is optimized by updating the latent gates using back propagation to minimize the distance between the evolved and target states. Taking the ground states of quantum lattice models and random matrix product states as examples, with the number of qubits where processing the full coefficients is unlikely, ADQC obtains high fidelities with small numbers of layers $N_L \sim O(1)$. Superior accuracy is reached compared with the existing state-preparation approach based on the matrix product disentangler. The parameter complexity of MPS can be significantly reduced by ADQC with the compression ratio $r \sim O(10^{-3})$. Our work sheds light on the "intelligent construction" of quantum circuits for many-qubit systems by combining with the machine learning methods.

quant-ph