Search arXivSearch

arXiv subjects

Chen Guo

Publications and source records attributed to Chen Guo.

At least 19 recordsLinked to original sources

Near single-cycle pulse generation using a cascaded all-bulk multi-pass cell compression system

We experimentally demonstrate the generation of sub-two-cycle optical pulses using a cascaded post-compression scheme based on two bulk multi-pass cells. Starting from 220 fs, 100 ${\mu}$J pulses around 1030 nm, sequential spectral broadening and pulse compression reduce the pulse duration first to 50 fs and ultimately to 5.5 fs (1.6 optical cycles) at a central wavelength of 1058 nm. The ultra-broadband spectrum is enabled by a dispersion-controlled cavity. Comprehensive spectral, temporal, and spatial characterization confirms excellent pulse quality. Despite operating the second multi-pass cell at peak powers several hundred times above the critical power for self-focusing in fused silica, no significant spatio-spectral or spatio-temporal distortions are observed, enabling direct use of the compressed 40 ${\mu}$J pulses in further experiments. Numerical simulations of the nonlinear spectral broadening show excellent agreement with the experimental results, supporting the underlying physical picture. These findings establish bulk multi-pass cells as an efficient, compact, and robust platform for generating few-cycle pulses with excellent beam quality, offering considerable potential for strong-field and ultrafast applications, including extreme-ultraviolet generation and isolated attosecond pulse production.

physics.optics

RHINO: Reconstructing Human Interactions with Novel Objects from Monocular Videos

Reconstructing people, objects, and their interactions in 3D is a long-standing goal for intelligent systems. Often the input is RGB video from a moving camera, making the task ill-posed; depth is ambiguous, humans and objects occlude each other, and camera and object motion entangle to create apparent motion. Most prior work addresses humans or objects in isolation, ignoring their interplay, or assumes known 3D shapes or cameras, which is impractical for real-world applications. We develop RHINO (Reconstructing Human Interactions with Novel Objects), a three-step framework that recovers in 3D a human, novel (unseen) manipulated object, and static scene in a common world frame from a monocular RGB video. First, we leverage 3D-aware foundation models to obtain cues that stabilize Structure-from-Motion (SfM) even for low-texture regions; this yields a coarse shape and apparent motion of a manipulated object from foreground pixels, and a coarse scene shape and camera motion from background pixels. Second, we estimate a human in the camera frame via an off-the-shelf method, and subtract the camera motion from apparent motion to extract the object motion; this registers the human, object, and coarse scene shapes into a common world frame. Third, we refine shapes using a compositional neural field with per-component signed-distance fields. The latter further enables differentiable contact priors that attract surfaces while penalizing interpenetration, improving the physical plausibility of the final reconstruction. For evaluation, we capture a new dataset of handheld monocular videos synchronized with a volumetric 4D capture stage, providing ground-truth shape and camera motion. RHINO outperforms state-of-the-art baselines on novel-view synthesis and 4D reconstruction. Ablations show that each stage contributes substantially. Code and data are available at https://lxxue.github.io/RHINO.

cs.CV

Do BERT Embeddings Encode Narrative Dimensions? A Token-Level Probing Analysis of Time, Space, Causality, and Character in Fiction

Narrative understanding requires multidimensional semantic structures. This study investigates whether BERT embeddings encode dimensions of fictional narrative semantics -- time, space, causality, and character. Using an LLM to accelerate annotation, we construct a token-level dataset labeled with these four narrative categories plus "others." A linear probe on BERT embeddings (94% accuracy) significantly outperforms a control probe on variance-matched random embeddings (47%), confirming that BERT encodes meaningful narrative information. With balanced class weighting, the probe achieves a macro-average recall of 0.83, with moderate success on rare categories such as causality (recall = 0.75) and space (recall = 0.66). However, confusion matrix analysis reveals "Boundary Leakage," where rare dimensions are systematically misclassified as "others." Clustering analysis shows that unsupervised clustering aligns near-randomly with predefined categories (ARI = 0.081), suggesting that narrative dimensions are encoded but not as discretely separable clusters. Future work includes a POS-only baseline to disentangle syntactic patterns from narrative encoding, expanded datasets, and layer-wise probing.

cs.CL

Enhanced dynamic range spatio-spectral metrology of few-cycle laser pulses

Accurate spatio-temporal and spatio-spectral metrology is critical to the characterization and use of ultra-short, high-power lasers. The emergence of few cycle pulses, with bandwidths of tens or hundreds of nanometers, poses a significant challenge to existing metrology techniques. This is due both to large discrepancies in the sensitivities of the measurements at different wavelengths and to variation in the spectral intensity at those wavelengths. In this paper, the authors propose spectral filtering and stitching of the measurements as a robust, simple solution that enhances the dynamic range of the measurements, allowing accurate few-cycle pulse reconstruction. This enhancement is demonstrated using INSIGHT -- the most commonly used spatio-spectral measurement device -- as well as using IMPALA and spatially resolved Fourier transform spectrometry.

physics.optics

Gaussian Wardrobe: Compositional 3D Gaussian Avatars for Free-Form Virtual Try-On

We introduce Gaussian Wardrobe, a novel framework to digitalize compositional 3D neural avatars from multi-view videos. Existing methods for 3D neural avatars typically treat the human body and clothing as an inseparable entity. However, this paradigm fails to capture the dynamics of complex free-form garments and limits the reuse of clothing across different individuals. To overcome these problems, we develop a novel, compositional 3D Gaussian representation to build avatars from multiple layers of free-form garments. The core of our method is decomposing neural avatars into bodies and layers of shape-agnostic neural garments. To achieve this, our framework learns to disentangle each garment layer from multi-view videos and canonicalizes it into a shape-independent space. In experiments, our method models photorealistic avatars with high-fidelity dynamics, achieving new state-of-the-art performance on novel pose synthesis benchmarks. In addition, we demonstrate that the learned compositional garments contribute to a versatile digital wardrobe, enabling a practical virtual try-on application where clothing can be freely transferred to new subjects. Project page: https://ait.ethz.ch/gaussianwardrobe

cs.CV

FlexAvatar: Flexible Large Reconstruction Model for Animatable Gaussian Head Avatars with Detailed Deformation

We present FlexAvatar, a flexible large reconstruction model for high-fidelity 3D head avatars with detailed dynamic deformation from single or sparse images, without requiring camera poses or expression labels. It leverages a transformer-based reconstruction model with structured head query tokens as canonical anchor to aggregate flexible input-number-agnostic, camera-pose-free and expression-free inputs into a robust canonical 3D representation. For detailed dynamic deformation, we introduce a lightweight UNet decoder conditioned on UV-space position maps, which can produce detailed expression-dependent deformations in real time. To better capture rare but critical expressions like wrinkles and bared teeth, we also adopt a data distribution adjustment strategy during training to balance the distribution of these expressions in the training set. Moreover, a lightweight 10-second refinement can further enhances identity-specific details in extreme identities without affecting deformation quality. Extensive experiments demonstrate that our FlexAvatar achieves superior 3D consistency, detailed dynamic realism compared with previous methods, providing a practical solution for animatable 3D avatar creation.

cs.CV

PHD: Personalized 3D Human Body Fitting with Point Diffusion

We introduce PHD, a novel approach for personalized 3D human mesh recovery (HMR) and body fitting that leverages user-specific shape information to improve pose estimation accuracy from videos. Traditional HMR methods are designed to be user-agnostic and optimized for generalization. While these methods often refine poses using constraints derived from the 2D image to improve alignment, this process compromises 3D accuracy by failing to jointly account for person-specific body shapes and the plausibility of 3D poses. In contrast, our pipeline decouples this process by first calibrating the user's body shape and then employing a personalized pose fitting process conditioned on that shape. To achieve this, we develop a body shape-conditioned 3D pose prior, implemented as a Point Diffusion Transformer, which iteratively guides the pose fitting via a Point Distillation Sampling loss. This learned 3D pose prior effectively mitigates errors arising from an over-reliance on 2D constraints. Consequently, our approach improves not only pelvis-aligned pose accuracy but also absolute pose accuracy -- an important metric often overlooked by prior work. Furthermore, our method is highly data-efficient, requiring only synthetic data for training, and serves as a versatile plug-and-play module that can be seamlessly integrated with existing 3D pose estimators to enhance their performance. Project page: https://phd-pose.github.io/

cs.CV

A New proof of Liouville type theorems for a class of semilinear elliptic equations

We study certain typical semilinear elliptic equations in Euclidean space $\bR^{n}$ or on a closed manifold $M$ with nonnegative Ricci curvature. Our proof is based on a crucial integral identity constructed by the invariant tensor method. Together with suitable integral estimates, some classical Liouville theorems will be reestablished.

math.AP

Compact, intense attosecond sources driven by hollow Gaussian beams

High-order harmonic generation (HHG) enables the up-conversion of intense infrared or visible femtosecond laser pulses into extreme-ultraviolet attosecond pulses. However, the highly nonlinear nature of the process results in low conversion efficiency, which can be a limitation for applications requiring substantial pulse energy, such as nonlinear attosecond time-resolved spectroscopy or single-shot diffractive imaging. Refocusing of the attosecond pulses is also essential to achieve a high intensity, but difficult in practice due to strong chromatic aberrations. In this work, we address both the generation and the refocusing of attosecond pulses by sculpting the driving beam into a ring-shaped intensity profile with no spatial phase variations, referred to as a Hollow Gaussian beam (HGB). Our experimental and theoretical results reveal that HGBs efficiently redistribute the driving laser energy in the focus, where the harmonics are generated on a ring with low divergence, which furthermore decreases with increasing order. Although generated as a ring, the attosecond pulses can be refocused with greatly reduced chromatic spread, therefore reaching higher intensity. This approach enhances the intensity of refocused attosecond pulses and enables significantly higher energy to be delivered in the driving beam without altering the focusing conditions. These combined advantages open pathways for compact, powerful, tabletop, laser-driven attosecond light sources.

physics.optics

SEGA: Drivable 3D Gaussian Head Avatar from a Single Image

Creating photorealistic 3D head avatars from limited input has become increasingly important for applications in virtual reality, telepresence, and digital entertainment. While recent advances like neural rendering and 3D Gaussian splatting have enabled high-quality digital human avatar creation and animation, most methods rely on multiple images or multi-view inputs, limiting their practicality for real-world use. In this paper, we propose SEGA, a novel approach for Single-imagE-based 3D drivable Gaussian head Avatar creation that combines generalized prior models with a new hierarchical UV-space Gaussian Splatting framework. SEGA seamlessly combines priors derived from large-scale 2D datasets with 3D priors learned from multi-view, multi-expression, and multi-ID data, achieving robust generalization to unseen identities while ensuring 3D consistency across novel viewpoints and expressions. We further present a hierarchical UV-space Gaussian Splatting framework that leverages FLAME-based structural priors and employs a dual-branch architecture to disentangle dynamic and static facial components effectively. The dynamic branch encodes expression-driven fine details, while the static branch focuses on expression-invariant regions, enabling efficient parameter inference and precomputation. This design maximizes the utility of limited 3D data and achieves real-time performance for animation and rendering. Additionally, SEGA performs person-specific fine-tuning to further enhance the fidelity and realism of the generated avatars. Experiments show our method outperforms state-of-the-art approaches in generalization ability, identity preservation, and expression realism, advancing one-shot avatar creation for practical applications.

cs.GR

Gradient Estimates for the doubly nonlinear diffusion equation on Complete Riemannian Manifolds

We study the elliptic version of doubly nonlinear diffusion equations on a complete Riemannian manifold $(M,g)$. Through the combination of a special nonlinear transformation and the standard Nash-Moser iteration procedure, some Cheng-Yau type gradient estimates for positive solutions are derived. As by-products, we also obtain Liouville type results and Harnack's inequality. These results fill a gap in Yan and Wang (2018)\cite{YW}, due to the lack of one key inequality when $b=\gamma-\frac{1}{p-1}>0$, and provide a partial answer to the question that whether gradient estimates for the doubly nonlinear diffusion equation can be extended to the case $b>0$ .

math.AP

Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal Prior

We present Vid2Avatar-Pro, a method to create photorealistic and animatable 3D human avatars from monocular in-the-wild videos. Building a high-quality avatar that supports animation with diverse poses from a monocular video is challenging because the observation of pose diversity and view points is inherently limited. The lack of pose variations typically leads to poor generalization to novel poses, and avatars can easily overfit to limited input view points, producing artifacts and distortions from other views. In this work, we address these limitations by leveraging a universal prior model (UPM) learned from a large corpus of multi-view clothed human performance capture data. We build our representation on top of expressive 3D Gaussians with canonical front and back maps shared across identities. Once the UPM is learned to accurately reproduce the large-scale multi-view human images, we fine-tune the model with an in-the-wild video via inverse rendering to obtain a personalized photorealistic human avatar that can be faithfully animated to novel human motions and rendered from novel views. The experiments show that our approach based on the learned universal prior sets a new state-of-the-art in monocular avatar reconstruction by substantially outperforming existing approaches relying only on heuristic regularization or a shape prior of minimally clothed bodies (e.g., SMPL) on publicly available datasets.

cs.CV

Influence of the laser pulse duration in high-order harmonic generation

High-order harmonic generation (HHG) in gases has been studied for almost 40 years in many different conditions, varying the laser wavelength, intensity, focusing geometry, target design, gas species, etc. However, no systematic investigation of the effect of the pulse duration has been performed in spite of its expected impact on phase-matching of the high-order harmonics. Here, we develop a compact post-compression method based on a bulk multi-pass cell enabling tunable Fourier-limited pulse durations. We examine the HHG yield as a function of the pulse duration, ranging from 42 fs to 180 fs, while maintaining identical focusing conditions and generating medium. Our findings reveal that, for a given intensity, there exists an optimum pulse duration - not necessarily the shortest - that maximizes conversion efficiency. This optimum pulse duration increases as the intensity decreases. The experimental results are corroborated by numerical simulations, which show the dependence of HHG yield on the duration and peak intensity of the driving laser and underscore the importance of the interplay between light-matter interaction and phase-matching in the non-linear medium. Our conclusion explains why HHG could be demonstrated in 1988 with pulses as long as 40 ps and intensities of just a few $10^{13}$ W/cm${^2}$.

physics.optics

Energy scaling in a compact bulk multi-pass cell enabled by Laguerre-Gaussian single-vortex beams

We report pulse energy scaling enabled by the use of Laguerre-Gaussian single-vortex ($\text{LG}_{0,l}$) beams for spectral broadening in a sub-40 cm long Herriott-type bulk multi-pass cell. Beams with orders ${l= 1-3}$ are generated by a spatial light modulator, which facilitates rapid and precise reconfiguration of the experimental conditions. 180 fs pulses with 610 uJ pulse energy are post-compressed to 44 fs using an $\text{LG}_{0,3}$ beam, boosting the peak power of an Ytterbium laser system from 2.5 GW to 9.1 GW. The spatial homogeneity of the output $\text{LG}_{0,l}$ beams is quantified and the topological charge is spectrally-resolved and shown to be conserved after compression by employing a custom spatio-temporal coupling measurement setup.

physics.optics

XUV yield optimization of two-color high-order harmonic generation in gases

We perform an experimental two-color high-order harmonic generation study in argon with the fundamental of an ytterbium ultrashort pulse laser and its second harmonic. The intensity of the second harmonic and its phase relative to the fundamental are varied, in a large range compared to earlier works, while keeping the total intensity constant. We extract the optimum values for the relative phase and ratio of the two colors which lead to a maximum yield enhancement for each harmonic order in the extreme ultraviolet spectrum. Within the semi-classical three-step model, the yield maximum can be associated with a flat electron return time vs. return energy distribution. An analysis of different distributions allows to predict the required relative two-color phase and ratio for a given harmonic order, total laser intensity, fundamental wavelength, and ionization potential.

physics.atom-ph

Multiscale carrier-envelope phase characterization of 2-\mu m pulses delivered by a 200-kHz optical parametric amplifier

Light fields with a central wavelength of 2 um are very well suited for strong-field-driven charge carrier control: Their photon energy lies far below the band gap of many materials, while their oscillation period remains significantly shorter than the coherence time of charge carrier oscillations. The resulting potential for field-driven charge carrier control is contingent on the reproducibility of the field structure of such ultrashort laser pulses. Here, we present a compact 200-kHz laser system that delivers ultrashort pulses with a duration of less than 20 fs in the spectral range around 2 um and with a pulse energy of 25 uJ. The electric field structure of the 2-um pulses is characterized in detail. In particular, the carrier-envelope phase (CEP) is measured over a wide range of timescales, from microseconds to hours. Passive stabilization due to difference frequency generation results in a root mean square value of carrier-envelope phase noise of less than 70 mrad over all measured time scales. The applicability of the pulses is demonstrated by measuring CEP-dependent high-order harmonic spectra with energies of up to 160 eV.

physics.optics

Packing Analysis: Packing Is More Appropriate for Large Models or Datasets in Supervised Fine-tuning

Packing, initially utilized in the pre-training phase, is an optimization technique designed to maximize hardware resource efficiency by combining different training sequences to fit the model's maximum input length. Although it has demonstrated effectiveness during pre-training, there remains a lack of comprehensive analysis for the supervised fine-tuning (SFT) stage on the following points: (1) whether packing can effectively enhance training efficiency while maintaining performance, (2) the suitable size of the model and dataset for fine-tuning with the packing method, and (3) whether packing unrelated or related training samples might cause the model to either excessively disregard or over-rely on the context. In this paper, we perform extensive comparisons between SFT methods using padding and packing, covering SFT datasets ranging from 69K to 1.2M and models from 8B to 70B. This provides the first comprehensive analysis of the advantages and limitations of packing versus padding, as well as practical considerations for implementing packing in various training scenarios. Our analysis covers various benchmarks, including knowledge, reasoning, and coding, as well as GPT-based evaluations, time efficiency, and other fine-tuning parameters. We also open-source our code for fine-tuning and evaluation and provide checkpoints fine-tuned on datasets of different sizes, aiming to advance future research on packing methods. Code is available at: https://github.com/ShuheWang1998/Packing-Analysis?tab=readme-ov-file.

cs.LG

ReLoo: Reconstructing Humans Dressed in Loose Garments from Monocular Video in the Wild

While previous years have seen great progress in the 3D reconstruction of humans from monocular videos, few of the state-of-the-art methods are able to handle loose garments that exhibit large non-rigid surface deformations during articulation. This limits the application of such methods to humans that are dressed in standard pants or T-shirts. Our method, ReLoo, overcomes this limitation and reconstructs high-quality 3D models of humans dressed in loose garments from monocular in-the-wild videos. To tackle this problem, we first establish a layered neural human representation that decomposes clothed humans into a neural inner body and outer clothing. On top of the layered neural representation, we further introduce a non-hierarchical virtual bone deformation module for the clothing layer that can freely move, which allows the accurate recovery of non-rigidly deforming loose clothing. A global optimization jointly optimizes the shape, appearance, and deformations of the human body and clothing via multi-layer differentiable volume rendering. To evaluate ReLoo, we record subjects with dynamically deforming garments in a multi-view capture studio. This evaluation, both on existing and our novel dataset, demonstrates ReLoo's clear superiority over prior art on both indoor datasets and in-the-wild videos.

cs.CV