Search arXivSearch

arXiv subjects

Zhenyu Xie

Publications and source records attributed to Zhenyu Xie.

At least 19 recordsLinked to original sources

TempAct: Advancing Temporal Plausibility in Autoregressive Video Generation via Planner-Executor RL

Autoregressive (AR) video diffusion models enable low-latency streaming generation by synthesizing videos chunk by chunk with cached visual context, but this chunk-wise formulation makes temporal instruction following ambiguous. A single global prompt does not specify which sub-event should be realized in each chunk, while naively switching to step-wise prompts often leads to delayed reactions, blended step semantics, and error propagation across prompt transitions. These failures are difficult to address with supervised fine-tuning or distillation alone: SFT suffers from exposure bias, while rollout-based distillation still optimizes low-level denoising or teacher-distribution matching rather than directly enforcing action ordering and prompt-transition correctness. We address these challenges with TempAct, a planner--executor reinforcement learning framework that jointly optimizes temporal decomposition and step-conditioned execution for temporally plausible AR video generation. TempAct uses an LLM planner to explore span-aware step prompts that are executable by the video model, and trains an AR diffusion executor to follow these prompts under its own generated histories. Its key mechanism is hierarchical group exploration: candidate plans form planning groups, and each plan induces an execution group of multiple continuations from a shared visual context, enabling plan-level credit assignment for long-horizon temporal outcomes and executor-level credit assignment for prompt-switch behavior. We further design hierarchical rewards that combine plan-quality and full-video temporal feedback for the planner with local transition-level step-following rewards, aesthetic regularization, and KL constraints for the executor. Experiments on Self-Forcing and LongLive show that TempAct improves temporal consistency while preserving overall visual quality.

cs.CV

Efficient Generative Retrieval for E-commerce Search with Semantic Cluster IDs and Expert-Guided RL

Generative retrieval offers a promising alternative by unifying the fragmented multi-stage retrieval process into a single end-to-end model. However, its practical adoption in industrial e-commerce search remains challenging, given the massive and dynamic product catalogs, strict latency requirements, and the need to align retrieval with downstream ranking goals. In this work, we propose a retrieval framework tailored for real-world recall scenarios, positioning generative retrieval as a recall-stage supplement rather than an end-to-end replacement. Our method, CQ-SID (Category-and-Query constrained Semantic ID), employs category-aware and query-item contrastive learning along with Residual Quantized VAEs to encode items into hierarchical semantic cluster identifiers, significantly reducing beam search complexity. Additionally, we develop EG-GRPO (Expert-Guided Group Relative Policy Optimization), a reinforcement learning approach that aligns generative recall with downstream ranking under sparse rewards by injecting ground-truth samples to stabilize training. Offline experiments on TmallAPP search logs show that CQ-SID achieves up to 26.76% and 11.11% relative gains in semantic and personalized click hitrate over RQ-VAE baselines, while halving beam search size. EG-GRPO further improves multi-objective performance. Online A/B tests confirm gains in GMV (+1.15%) and UCTCVR (+0.40%). The generative recall channel now contributes substantially in production, accounting for over 50.25% of exposures, 58.96% of clicks, and 72.63% of purchases, demonstrating a viable path for deploying generative retrieval in real-world e-commerce systems.

cs.IR

HairOrbit: Multi-view Aware 3D Hair Modeling from Single Portraits

Reconstructing strand-level 3D hair from a single-view image is highly challenging, especially when preserving consistent and realistic attributes in unseen regions. Existing methods rely on limited frontal-view cues and small-scale/style-restricted synthetic data, often failing to produce satisfactory results in invisible regions. In this work, we propose a novel framework that leverages the strong 3D priors of video generation models to transform single-view hair reconstruction into a calibrated multi-view reconstruction task. To balance reconstruction quality and efficiency for the reformulated multi-view task, we further introduce a neural orientation extractor trained on sparse real-image annotations for better full-view orientation estimation. In addition, we design a two-stage strand-growing algorithm based on a hybrid implicit field to synthesize the 3D strand curves with fine-grained details at a relatively fast speed. Extensive experiments demonstrate that our method achieves state-of-the-art performance on single-view 3D hair strand reconstruction on a diverse range of hair portraits in both visible and invisible regions.

cs.CV

Steering Video Diffusion Transformers with Massive Activations

In this work, we study the role of Massive Activations (MAs), which are rare, high-magnitude spikes confined to a few fixed hidden dimensions in video diffusion transformers (DiTs). We uncover a structured positional hierarchy: MA magnitudes peak at first-frame tokens and recur at the spatial boundary tokens of latent frames, with this pattern being most pronounced during early denoising. We trace this organization to an encoding asymmetry of the video VAEs, whose causal temporal padding and zero spatial padding cause the first latent frame and frame borders to carry reduced content load. Elevated MAs consistently align with these lower-content structural positions. To understand their function, we analyze intermediate representations and find that MAs act as implicit rescalers of residual computation: enlarging MAs suppresses the corresponding self-attention and feed-forward updates, while erasing them amplifies these updates. Together, these observations suggest that MAs serves as a token-level rescaler of residual computation, which video DiTs deploy unevenly, placing the strongest damping at the encoding-asymmetric structural positions. Motivated by this native rescaling behavior, we propose Structured Activation Steering (STAS), a training-free technique that steers MAs at the observed structural positions toward a scaled, model-derived reference during early denoising. STAS requires no additional forward passes and consistently improves video quality and temporal coherence across text-to-video models with negligible overhead.

cs.CV

AnyCrowd: Instance-Isolated Identity-Pose Binding for Arbitrary Multi-Character Animation

Controllable character animation has advanced rapidly in recent years, yet multi-character animation remains underexplored. As the number of characters grows, multi-character reference encoding becomes more susceptible to latent identity entanglement, resulting in identity bleeding and reduced controllability. Moreover, learning precise and spatio-temporally consistent correspondences between reference identities and driving pose sequences becomes increasingly challenging, often leading to identity-pose mis-binding and inconsistency in generated videos. To address these challenges, we propose AnyCrowd, a Diffusion Transformer (DiT)-based video generation framework capable of scaling to an arbitrary number of characters. Specifically, we first introduce an Instance-Isolated Latent Representation (IILR), which encodes character instances independently prior to DiT processing to prevent latent identity entanglement. Building on this disentangled representation, we further propose Tri-Stage Decoupled Attention (TSDA) to bind identities to driving poses by decomposing self-attention into: (i) instance-aware foreground attention, (ii) background-centric interaction, and (iii) global foreground-background coordination. Furthermore, to mitigate token ambiguity in overlapping regions, an Adaptive Gated Fusion (AGF) module is integrated within TSDA to predict identity-aware weights, effectively fusing competing token groups into identity-consistent representations...

cs.CV

ERGO: Excess-Risk-Guided Optimization for High-Fidelity Monocular 3D Gaussian Splatting

Generating 3D content from a single image remains a fundamentally challenging and ill-posed problem due to the inherent absence of geometric and textural information in occluded regions. While state-of-the-art generative models can synthesize auxiliary views to provide additional supervision, these views inevitably contain geometric inconsistencies and textural misalignments that propagate and amplify artifacts during 3D reconstruction. To effectively harness these imperfect supervisory signals, we propose an adaptive optimization framework guided by excess risk decomposition, termed ERGO. Specifically, ERGO decomposes the optimization losses in 3D Gaussian splatting into two components, i.e., excess risk that quantifies the suboptimality gap between current and optimal parameters, and Bayes error that models the irreducible noise inherent in synthesized views. This decomposition enables ERGO to dynamically estimate the view-specific excess risk and adaptively adjust loss weights during optimization. Furthermore, we introduce geometry-aware and texture-aware objectives that complement the excess-risk-derived weighting mechanism, establishing a synergistic global-local optimization paradigm. Consequently, ERGO demonstrates robustness against supervision noise while consistently enhancing both geometric fidelity and textural quality of the reconstructed 3D content. Extensive experiments on the Google Scanned Objects dataset and the OmniObject3D dataset demonstrate the superiority of ERGO over existing state-of-the-art methods.

cs.CV

Hertz-Integral-Linewidth Lasers based on Portable Solid-state Microresonators

Optical reference resonators serve as a cornerstone in various scientific fields. In recent years, there has been an increasing demand for compact ultrastable reference resonators capable of operating in ambient environments, enabling applications beyond the laboratory, such as navigation, portable optical clocks, and remote sensing. Here, we present a compact ultrastable whispering-gallery-mode \ce{MgF2} reference resonator with a high loaded quality factor of $2.24\times 10^9$. The device is packaged in a compact form of 50$\times$77$\times$90 mm and supports stable optical coupling with polarization-maintaining fiber, which enables robust operation under ambient conditions. Laser stabilization using this resonator yields a phase noise of -105 dBc/Hz at a 10 kHz offset frequency, an integral linewidth of 4 Hz, and a fractional frequency stability of $2.5\times 10^{-14}$ at a 10 ms averaging time. With the high performance and rapid manufacturability, our work offers a promising solution for ultrastable optical frequency references beyond laboratory settings.

physics.optics

Communication-ready high-power soliton microcombs in highly-dispersive Fabry-Perot-microresonators

Microcombs generated in optical microresonators are widely regarded as promising light sources for next-generation communication systems, but the optical power available per comb line has so far fallen short of practical requirements. Here we introduce an integrated Fabry-P\'erot microresonator platform that overcomes fundamental dispersion-engineering constraints and enables bright soliton microcombs with unprecedented power per line. The resonator is defined by chirped Bragg gratings that provide exceptionally large anomalous group-velocity dispersion, allowing more than ten comb lines to reach the milliwatt level. These combs can be used directly in coherent communication systems without additional amplification, achieving an aggregate data rate of 2 Tb/s. Once integrated, our high-power soliton microcombs could be instantly ready for communications as well as a broad range of practical comb-based applications.

physics.optics

GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping

Recently, GRPO-based reinforcement learning has shown remarkable progress in optimizing flow-matching models, effectively improving their alignment with task-specific rewards. Within these frameworks, the policy update relies on importance-ratio clipping to constrain overconfident positive and negative gradients. However, in practice, we observe a systematic shift in the importance-ratio distribution-its mean falls below 1 and its variance differs substantially across timesteps. This left-shifted and inconsistent distribution prevents positive-advantage samples from entering the clipped region, causing the mechanism to fail in constraining overconfident positive updates. As a result, the policy model inevitably enters an implicit over-optimization stage-while the proxy reward continues to increase, essential metrics such as image quality and text-prompt alignment deteriorate sharply, ultimately making the learned policy impractical for real-world use. To address this issue, we introduce GRPO-Guard, a simple yet effective enhancement to existing GRPO frameworks. Our method incorporates ratio normalization, which restores a balanced and step-consistent importance ratio, ensuring that PPO clipping properly constrains harmful updates across denoising timesteps. In addition, a gradient reweighting strategy equalizes policy gradients over noise conditions, preventing excessive updates from particular timestep regions. Together, these designs act as a regulated clipping mechanism, stabilizing optimization and substantially mitigating implicit over-optimization without relying on heavy KL regularization. Extensive experiments on multiple diffusion backbones (e.g., SD3.5M, Flux.1-dev) and diverse proxy tasks demonstrate that GRPO-Guard significantly reduces over-optimization while maintaining or even improving generation quality.

cs.CV

Power-efficient ultra-broadband soliton microcombs in resonantly-coupled microresonators

The drive to miniaturize optical frequency combs for practical deployment has spotlighted microresonator solitons as a promising chip-scale candidate. However, these soliton microcombs could be very power-hungry when their span increases, especially with fine comb spacings. As a result, realizing an octave-spanning comb at microwave repetition rates for direct optical-microwave linkage is considered not possible for photonic integration due to the high power requirements. Here, we introduce the concept of resonant-coupling to soliton microcombs to reduce pump consumption significantly. Compared to conventional waveguide-coupled designs, we demonstrate (i) a threefold increase in spectral span for high-power combs and (ii) up to a tenfold reduction in repetition frequency for octave-spanning operation. This configuration is compatible with laser integration and yields reliable, turnkey soliton generation. By eliminating the long-standing pump-power bottleneck, microcombs will soon become readily available for portable optical clocks, massively parallel data links, and field-deployable spectrometers.

physics.optics

Compact Turnkey Soliton Microcombs at Microwave Rates via Wafer-Scale Fabrication

Soliton microcombs generated in nonlinear microresonators facilitate the photonic integration of timing, frequency synthesis, and astronomical calibration functionalities. For these applications, low-repetition-rate soliton microcombs are essential as they establish a coherent link between optical and microwave signals. However, the required pump power typically scales with the inverse of the repetition rate, and the device footprint scales with the inverse of square of the repetition rate, rendering low-repetition-rate soliton microcombs challenging to integrate within photonic circuits. This study designs and fabricates silicon nitride microresonators on 4-inch wafers with highly compact form factors. The resonator geometries are engineered from ring to finger and spiral shapes to enhance integration density while attaining quality factors over 10^7. Driven directly by an integrated laser, soliton microcombs with repetition rates below 10 GHz are demonstrated via turnkey initiation. The phase noise performance of the synthesized microwave signals reaches -130 dBc/Hz at 100 kHz offset frequency for 10 GHz carrier frequencies. This work enables the high-density integration of soliton microcombs for chip-based microwave photonics and spectroscopy applications.

physics.optics

Soliton microcombs in X-cut LiNbO3 microresonators

Chip-scale integration of optical frequency combs, particularly soliton microcombs, enables miniaturized instrumentation for timekeeping, ranging, and spectroscopy. Although soliton microcombs have been demonstrated on various material platforms, realizing complete comb functionality on photonic chips requires the co-integration of high-speed modulators and efficient frequency doublers, features that are available in a monolithic form on X-cut thin-film lithium niobate (TFLN). However, the pronounced Raman nonlinearity associated with extraordinary light in this platform has so far precluded soliton microcomb generation. Here, we report the generation of transverse-electric-polarized soliton microcombs with a 25 GHz repetition rate in high-Q microresonators on X-cut TFLN chips. By precisely orienting the racetrack microresonator relative to the optical axis, we mitigate Raman nonlinearity and enable soliton formation under continuous-wave laser pumping. Moreover, the soliton microcomb spectra are extended to 350 nm with pulsed laser pumping. This work expands the capabilities of TFLN photonics and paves the way for the monolithic integration of fast-tunable, self-referenced microcombs.

physics.optics

DreamVTON: Customizing 3D Virtual Try-on with Personalized Diffusion Models

Image-based 3D Virtual Try-ON (VTON) aims to sculpt the 3D human according to person and clothes images, which is data-efficient (i.e., getting rid of expensive 3D data) but challenging. Recent text-to-3D methods achieve remarkable improvement in high-fidelity 3D human generation, demonstrating its potential for 3D virtual try-on. Inspired by the impressive success of personalized diffusion models (e.g., Dreambooth and LoRA) for 2D VTON, it is straightforward to achieve 3D VTON by integrating the personalization technique into the diffusion-based text-to-3D framework. However, employing the personalized module in a pre-trained diffusion model (e.g., StableDiffusion (SD)) would degrade the model's capability for multi-view or multi-domain synthesis, which is detrimental to the geometry and texture optimization guided by Score Distillation Sampling (SDS) loss. In this work, we propose a novel customizing 3D human try-on model, named \textbf{DreamVTON}, to separately optimize the geometry and texture of the 3D human. Specifically, a personalized SD with multi-concept LoRA is proposed to provide the generative prior about the specific person and clothes, while a Densepose-guided ControlNet is exploited to guarantee consistent prior about body pose across various camera views. Besides, to avoid the inconsistent multi-view priors from the personalized SD dominating the optimization, DreamVTON introduces a template-based optimization mechanism, which employs mask templates for geometry shape learning and normal/RGB templates for geometry/texture details learning. Furthermore, for the geometry optimization phase, DreamVTON integrates a normal-style LoRA into personalized SD to enhance normal map generative prior, facilitating smooth geometry modeling.

cs.CV

MMTryon: Multi-Modal Multi-Reference Control for High-Quality Fashion Generation

This paper introduces MMTryon, a multi-modal multi-reference VIrtual Try-ON (VITON) framework, which can generate high-quality compositional try-on results by taking a text instruction and multiple garment images as inputs. Our MMTryon addresses three problems overlooked in prior literature: 1) \textbf{Support of multiple try-on items.} Existing methods are commonly designed for single-item try-on tasks (e.g., upper/lower garments, dresses). 2) \textbf{Specification of dressing style}. Existing methods are unable to customize dressing styles based on instructions (e.g., zipped/unzipped, tuck-in/tuck-out, etc.) 3) \textbf{Segmentation Dependency}. They further heavily rely on category-specific segmentation models to identify the replacement regions, with segmentation errors directly leading to significant artifacts in the try-on results. To address the first two issues, our MMTryon introduces a novel multi-modality and multi-reference attention mechanism to combine the garment information from reference images and dressing-style information from text instructions. Besides, to remove the segmentation dependency, MMTryon uses a parsing-free garment encoder and leverages a novel scalable data generation pipeline to convert existing VITON datasets to a form that allows MMTryon to be trained without requiring any explicit segmentation. Extensive experiments on high-resolution benchmarks and in-the-wild test sets demonstrate MMTryon's superiority over existing SOTA methods both qualitatively and quantitatively. MMTryon's impressive performance on multi-item and style-controllable virtual try-on scenarios and its ability to try on any outfit in a large variety of scenarios from any source image, opens up a new avenue for future investigation in the fashion community.

cs.CV

Microresonator-referenced soliton microcombs with zeptosecond-level timing noise

Optical frequency division relies on optical frequency combs to coherently translate ultra-stable optical frequency references to the microwave domain. This technology has enabled microwave synthesis with ultralow timing noise, but the required instruments are too bulky for real-world applications. Here, we develop a compact optical frequency division system using microresonator-based frequency references and comb generators. The soliton microcomb formed in an integrated Si$_3$N$_4$ microresonator is stabilized to two lasers referenced to an ultrahigh-$Q$ MgF$_2$ microresonator. Photodetection of the soliton pulse train produces 25 GHz microwaves with absolute phase noise of -141 dBc/Hz (547 zs Hz$^{-1/2}$) at 10 kHz offset frequency. The synthesized microwaves are tested as local oscillators in jammed communication channels, resulting in improved fidelity compared with those derived from electronic oscillators. Our work demonstrates unprecedented coherence in miniature microwave oscillators, providing key building blocks for next-generation timekeeping, navigation, and satellite communication systems.

physics.optics

GUESS:GradUally Enriching SyntheSis for Text-Driven Human Motion Generation

In this paper, we propose a novel cascaded diffusion-based generative framework for text-driven human motion synthesis, which exploits a strategy named GradUally Enriching SyntheSis (GUESS as its abbreviation). The strategy sets up generation objectives by grouping body joints of detailed skeletons in close semantic proximity together and then replacing each of such joint group with a single body-part node. Such an operation recursively abstracts a human pose to coarser and coarser skeletons at multiple granularity levels. With gradually increasing the abstraction level, human motion becomes more and more concise and stable, significantly benefiting the cross-modal motion synthesis task. The whole text-driven human motion synthesis problem is then divided into multiple abstraction levels and solved with a multi-stage generation framework with a cascaded latent diffusion model: an initial generator first generates the coarsest human motion guess from a given text description; then, a series of successive generators gradually enrich the motion details based on the textual description and the previous synthesized results. Notably, we further integrate GUESS with the proposed dynamic multi-condition fusion mechanism to dynamically balance the cooperative effects of the given textual condition and synthesized coarse motion prompt in different generation stages. Extensive experiments on large-scale datasets verify that GUESS outperforms existing state-of-the-art methods by large margins in terms of accuracy, realisticness, and diversity. Code is available at https://github.com/Xuehao-Gao/GUESS.

cs.CV

Towards Detailed Text-to-Motion Synthesis via Basic-to-Advanced Hierarchical Diffusion Model

Text-guided motion synthesis aims to generate 3D human motion that not only precisely reflects the textual description but reveals the motion details as much as possible. Pioneering methods explore the diffusion model for text-to-motion synthesis and obtain significant superiority. However, these methods conduct diffusion processes either on the raw data distribution or the low-dimensional latent space, which typically suffer from the problem of modality inconsistency or detail-scarce. To tackle this problem, we propose a novel Basic-to-Advanced Hierarchical Diffusion Model, named B2A-HDM, to collaboratively exploit low-dimensional and high-dimensional diffusion models for high quality detailed motion synthesis. Specifically, the basic diffusion model in low-dimensional latent space provides the intermediate denoising result that to be consistent with the textual description, while the advanced diffusion model in high-dimensional latent space focuses on the following detail-enhancing denoising process. Besides, we introduce a multi-denoiser framework for the advanced diffusion model to ease the learning of high-dimensional model and fully explore the generative potential of the diffusion model. Quantitative and qualitative experiment results on two text-to-motion benchmarks (HumanML3D and KIT-ML) demonstrate that B2A-HDM can outperform existing state-of-the-art methods in terms of fidelity, modality consistency, and diversity.

cs.CV

WarpDiffusion: Efficient Diffusion Model for High-Fidelity Virtual Try-on

Image-based Virtual Try-On (VITON) aims to transfer an in-shop garment image onto a target person. While existing methods focus on warping the garment to fit the body pose, they often overlook the synthesis quality around the garment-skin boundary and realistic effects like wrinkles and shadows on the warped garments. These limitations greatly reduce the realism of the generated results and hinder the practical application of VITON techniques. Leveraging the notable success of diffusion-based models in cross-modal image synthesis, some recent diffusion-based methods have ventured to tackle this issue. However, they tend to either consume a significant amount of training resources or struggle to achieve realistic try-on effects and retain garment details. For efficient and high-fidelity VITON, we propose WarpDiffusion, which bridges the warping-based and diffusion-based paradigms via a novel informative and local garment feature attention mechanism. Specifically, WarpDiffusion incorporates local texture attention to reduce resource consumption and uses a novel auto-mask module that effectively retains only the critical areas of the warped garment while disregarding unrealistic or erroneous portions. Notably, WarpDiffusion can be integrated as a plug-and-play component into existing VITON methodologies, elevating their synthesis quality. Extensive experiments on high-resolution VITON benchmarks and an in-the-wild test set demonstrate the superiority of WarpDiffusion, surpassing state-of-the-art methods both qualitatively and quantitatively.

cs.CV