Search arXivSearch

arXiv subjects

Han Yan

Publications and source records attributed to Han Yan.

At least 19 recordsLinked to original sources

Fashion-3DLR: A Controllable 3D Garment Generation Using Pairwise Fashion Elements for Intelligent Design

AI-generated content (AIGC) has made significant progress, with 2D generative models becoming ready-to-use tools for the digital fashion industry. However, 3D garment generation remains in its nascent stage, where in the realm of fashion, the semantic information of diverse design elements exhibits intricate coupling relationships in 3D representations, posing substantial challenges for generating diverse 3D garments. In this work, to handle the above problem, We introduce Fashion-3DLR, a novel 3D garment generation framework that utilizes diverse design elements to create high-quality, versatile 3D garment assets. Specifically, to bridge the semantic gaps between different fashion elements, we propose a Garment Feature Fusion Diffusion Transformer (GFF-DiT) module to integrate 2D fashion design elements, e.g., sketch and texture, into latent space. Within the latent space, we then employ a rectified flow transformer to generate geometry latents, which can be decoded into various 3D garment representations, including 3D Gaussians and meshes. Furthermore, we integrate Fashion-3DLR into downstream tasks, achieving the 3D Gaussian Splatting (3DGS)-driven cloth physical simulation and mesh-based virtual try-on. Experimental results indicate that Fashion-3DLR surpass the previous state-of-the-art methods, which verify that the proposed work can generate well-structured, non-watertight garments capable of physical simulation and virtual try-on, underscoring its potential as a versatile 3D garment design tool.

cs.CV

Symmetry-Protected Pinch Curves in Classical Spin Liquids

Classical spin liquids are correlated paramagnets in which local constraints generate extensive degeneracy and emergent gauge structures, often observable as pinch-point singularities in spin structure factors. Here we introduce pinch-curve spin liquids, in which the pinch singularities form one-dimensional algebraic curves in momentum space. Inversion symmetry protects these curves by reducing the singularity condition to two real algebraic constraints in three dimensions, and the geometry of the pinch locus is algebraically programmable. We identify elementary mechanisms for generating straight and curved pinch loci, construct lattice spin models that realize them, and test the predicted structure factors using Monte Carlo simulations. We further show that pinch curves can host an infrared Gauss-law transition: the leading local constraint and the associated anisotropic scaling of the structure factor change, even though the singular locus remains one-dimensional.

cond-mat.str-el

Random Local Stabilizer Codes in Three Dimensions without String or Self-Similar Fractal Logical Operators

Quantum error-correcting codes (QECs) are essential components of quantum computation and have deep connections to quantum phases of matter. A key obstruction to passive self-correcting QECs is the presence of string logical operators, which can generate logical errors through constant-energy-barrier processes. Haah's Codes (fracton codes) showed that three-dimensional stabilizer codes can forbid such string logical operators, but their translation-invariant structure supports self-similar fractal logical operators with a logarithmic energy barrier. We introduce the qutrit random cubic codes, a family of local qutrit Calderbank-Shor-Steane stabilizer Hamiltonians with similar cube-check structure as Haah's Code 1 but built from spatially varying stabilizers. We prove that these models retain the no-string property and numerically observe that they have properties distinct from translation-invariant fracton codes: the smallest ground-state degeneracy exponent is $k=2$ for odd $L$ and $k=4$ for even $L$; noncontractible plane-logical operators span the entire logical space; and charge-push diagnostics show that the self-similar fractal operators are absent. These results demonstrate that constrained randomness can fundamentally change the nature of stabilizer codes and improve their self-correction properties. They further point to broader families of quantum error-correcting codes and quantum phases beyond canonical topological and fracton orders.

quant-ph

HiMed: Incentivizing Hindi Reasoning in Medical LLMs

Medical large language models hold promise for reducing healthcare disparities, yet Hindi remains severely underrepresented. While medical LLMs excel in high-resource languages, their performance degrades sharply in Hindi, particularly on Indian systems of medicine. We argue that robust cross-lingual medical transfer requires Hindi reasoning. To this end, we introduce HiMed, a Hindi reasoning medical corpus and benchmark suite covering both Western and Indian medicine. We further propose HiMed-8B, a Hindi-form medical reasoning LLM, through the design of decaying scaffolding reward. Extensive experiments demonstrate improvement in Hindi medical reasoning performance and reduction in the English--Hindi accuracy gap. Ablation studies validate the contribution of each training stage and reward component. All data and code are available on GitHub: https://github.com/FreedomIntelligence/HiMed.

cs.CL

Gravitational-wave Tomography of the Moon: Constraining Lunar Structure with Calibrated Gravitational Waves

The recent success of gravitational-wave (GW) astronomy together with renewed plans for lunar geophysical instrumentation has revived interest in using the Moon as a resonant detector for mid-frequency (mHz-Hz) GWs. In realistic observational scenarios, the GW strain amplitude is expected to be constrained independently by networks of GW detectors, which motivates an inverse, \emph{tomographic} question: to what extent can measurements of the Moon's seismic response to known GWs be used to infer its internal structure? In this work, we develop a first-principles, perturbative framework that maps spherically symmetric perturbations of the elastic and density structure to measurable changes in observables, especially GW-driven modal amplitudes of the Moon. The formalism combines (i) a normal-mode representation of the elastic response, (ii) first-order perturbation theory for eigenvalues and eigenfunctions, and (iii) a linearized observation model that links frequency and amplitude observables to model parameters (bulk and shear moduli, density, and interface locations) and their perturbations. We show that the estimation errors of the Moon's elastic parameters can be reduced by about an order of magnitude with observations of calibrated GWs.

astro-ph.EP

Seismic background mitigation with the Lunar Gravitational-wave Antenna

Lunar gravitational-wave (GW) detectors relying on the measurement of the response of the Moon to GWs are susceptible to a seismic background, which might pose a fundamental sensitivity limitation. The Lunar Gravitational-wave Antenna (LGWA) was conceived as an array of accelerometers with the idea that data can be processed to distinguish between a GW signal and the seismic background. As a result, the seismic noise of the GW measurement would be mitigated. However, so far, no quantitative assessment of the mitigation of the seismic background has been provided. In this article, we derive the analytical expressions for the optimal squared signal-to-noise ratio considering two seismic stations in an isotropic, random, Gaussian seismic field. Our numerical analysis reveals that the capacity to mitigate the seismic noise critically depends on the distance between the two stations relative to the seismic-correlation length. We demonstrate that optimal placement of the two stations can yield significant improvements in the equivalent seismic noise amplitude spectrum density (ASD), approximately a factor of 2.3 at 0.3 Hz, compared to the measurement with a single station. The equivalent ASD of the seismic noise also exhibits distinct oscillatory and mitigation features arising from the Bessel-function structure of the noise correlation.

astro-ph.IM

A Novel Method to Construct Frequency-Domain Gravitational Waveform for Accelerating Sources

Accurately modeling the inspiral-merger-ringdown (IMR) signal of coalescing compact objects is essential for the test of general relativity. However, it is known that astrophysical environments can distort gravitational-wave (GW) signal and, if ignored, may bias parameter estimation or even our understanding of gravity. Previous studies suggest that various astrophysical environmental effects can be modeled in a unified way by introducing an effective acceleration. However, such models are based on stationary phase approximation (SPA) and post-Newtonian (PN) formalism, which are inconsistent with the fast orbital evolution and strong gravity in the final merger-ringdown phase. To overcome this limit, we introduce frequency-domain spectral differentiation (FSD), which maps the time shift of the signal caused by acceleration into a differentiation in the frequency domain. The mapping does not rely on SPA or PN formalism, therefore can be used to construct the accelerated waveform across the entire IMR phases. We compare the FSD waveforms with the conventional SPA+PN ones, and find that the former more faithfully match the simulated signals of accelerating sources, especially in the merger-ringdown phase and when higher-order FSD corrections are included. A Fisher information matrix analysis suggests that FSD waveforms can achieve higher precision than SPA+PN waveforms in measuring effective acceleration. Therefore, the FSD method offers a more self-consistent treatment of various astrophysical environmental effects in the final merger-ringdown phase of binary GW sources.

astro-ph.HE

I3DM: Implicit 3D-aware Memory Retrieval and Injection for Consistent Video Scene Generation

Despite remarkable progress in video generation, maintaining long-term scene consistency upon revisiting previously explored areas remains challenging. Existing solutions rely either on explicitly constructing 3D geometry, which suffers from error accumulation and scale ambiguity, or on naive camera Field-of-View (FoV) retrieval, which typically fails under complex occlusions. To overcome these limitations, we propose I3DM, a novel implicit 3D-aware memory mechanism for consistent video scene generation that bypasses explicit 3D reconstruction. At the core of our approach is a 3D-aware memory retrieval strategy, which leverages the intermediate features of a pre-trained Feed-Forward Novel View Synthesis (FF-NVS) model to score view relevance, enabling robust retrieval even in highly occluded scenarios. Furthermore, to fully utilize the retrieved historical frames, we introduce a 3D-aligned memory injection module. This module implicitly warps historical content to the target view and adaptively conditions the generation on reliable warping regions, leading to improved revisit consistency and accurate camera control. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches, achieving superior revisit consistency, generation fidelity, and camera control precision.

cs.CV

MWM: Mobile World Models for Action-Conditioned Consistent Prediction

World models enable planning in imagined future predicted space, offering a promising framework for embodied navigation. However, existing navigation world models often lack action-conditioned consistency, so visually plausible predictions can still drift under multi-step rollout and degrade planning. Moreover, efficient deployment requires few-step diffusion inference, but existing distillation methods do not explicitly preserve rollout consistency, creating a training-inference mismatch. To address these challenges, we propose MWM, a mobile world model for planning-based image-goal navigation. Specifically, we introduce a two-stage training framework that combines structure pretraining with Action-Conditioned Consistency (ACC) post-training to improve action-conditioned rollout consistency. We further introduce Inference-Consistent State Distillation (ICSD) for few-step diffusion distillation with improved rollout consistency. Our experiments on benchmark and real-world tasks demonstrate consistent gains in visual fidelity, trajectory accuracy, planning success, and inference efficiency. Code: https://github.com/AIGeeksGroup/MWM. Website: https://aigeeksgroup.github.io/MWM.

cs.CV

Observation of Orbit-Orbit Torques: Highly Efficient Torques on Orbital Moments Induced by Orbital Currents

We study the current-induced torques in bilayers composed of a light 3d metal, chromium, and a rare-earth ferromagnet with finite orbital moments, terbium, utilizing second-harmonic Hall-response measurements. The dampinglike torque efficiency of chromium is found to be positive and reaches ~3.66 in this system, in sharp contrast to the negative and subtle dampinglike torque efficiency in general Cr/ferromagnet heterostructures with quenched orbital moment. We suggest that the orbital currents generated by the orbital Hall effect in Cr can be injected into Tb with negligible loss at the interface and then efficiently interact with the orbital moments. We term such an exotic effect as the orbit-orbit torque (OOT). Our work implies that orbital currents could be harnessed to manipulate the orbital magnetization of materials, which would advance the development of orbitronics.

cond-mat.mes-hall

Thick Lunar Crust Amplifies Deci-Hertz Gravitational-Wave Signal

Gravitational waves (GWs) in the $0.01\sim1$ Hz band encode unique signatures of the early universe and merging compact objects, but they are beyond the reach of existing observatories. Theoretical models suggest that the Moon could act as a resonant detector, but the unknown influence of its rugged surface and heterogeneous interior poses a challenge to the accurate modeling of its response. Here, we address this long-standing uncertainty by constructing the first high-resolution, two-dimensional model of the lunar GW response, more realistic than previous ones. We achieve this by combining high-fidelity spectral-element simulations with the analytical power of normal-mode perturbation theory, thereby resolving topographical effects down to 2 km grid spacing while maintaining the capacity to discern global free-oscillation patterns. This dual-methodology approach not only recovers the expected predominant quadrupole ($l=2$) oscillation mode, but also exposes a systematic signal amplification in thick-crust regions. This enhancement is traced by our normal-mode analysis to a mode-coupling process, in which the original quadrupolar oscillation induced by the passing GW distributes energy into a series of higher-order modes, the hybridized eigenmodes of a laterally heterogeneous Moon. In certain narrow frequency ranges, we observe up to tenfold amplification spanning into the deci-hertz band, highlighting the power of numerical simulations in resolving these structurally fine-tuned features for designing future detectors. Our work establishes the Moon as a resonant GW detector albeit its complex topographical structures, and the resulting amplification maps provide quantitative guide for the optimal landing site selection.

gr-qc

Numerical simulation of lunar response to gravitational waves and its 3D topographic effect using the spectral-element method

The Moon has been regarded as a natural Weber bar capable of amplifying gravitational waves (GWs) for detecting events across a wide range of frequencies. However, accurately determining the amplification effects remains challenging due to the absence of 3D numerical simulation methods. In this study, we develop a high-order 3D finite element method (spectral-element method, SEM) to numerically simulate the lunar response to GWs below 20 mHz. We verify the accuracy of our method by comparing the resonant peaks of our results with those from semi-analytical solutions and find that the frequency deviation is less than 3% for the first peak at about 1 mHz and less than 0.8% for the subsequent peaks up to 10 mHz. Using this method, we evaluate the amplification of GW signals due to 3D topographic effects of the Moon, and we find enhancements at a series of specific frequency components. These results highlight the non-negligible effect of surface topography on the lunar response to GWs, as a fundamental factor that holds significant implications across both global and regional analyses. Our work paves the way for a comprehensive evaluation of the Moon's resonant response to GWs, helpful for the strategic planning of lunar GW detections.

astro-ph.EP

Dielectric and gate metal engineering for threshold voltage modulation in enhancement mode monolayer MoS2 field effect transistors

Excellent gate electrostatics in field effect transistors (FETs) based on two-dimensional transition metal dichalcogenide (2D TMD) channels can dramatically decrease static power dissipation. Energy efficient FETs operate in enhancement mode with small and positive threshold voltage (Vth) for n-type devices. However, most state-of-the-art FETs based on monolayer MoS2 channel operate in depletion mode with negative Vth due to doping from the underlying dielectric substrate. In this work, we identify key properties of the semiconductor/dielectric interface (MoS2 on industrially relevant high dielectric constant (k) HfO2, ZrO2 and hBN for reference) responsible for realizing enhancement-mode operation of 2D MoS2 channel FETs. We find that hBN and ZrO2 dielectric substrates provide low defect interfaces with MoS2 that enables effective modulation of the Vth using gate metals of different work functions (WFs). We use photoluminescence (PL) and synchrotron X-ray photoelectron spectroscopy (XPS) measurements to investigate doping levels in monolayer MoS2 on different dielectrics with different WF gate metals. We complement the FET and spectroscopic measurements with capacitance-voltage analysis on dielectrics with varying thicknesses, which confirm that Vth modulation in ZrO2 devices is correlated with WF of the gate metals - in contrast with HfO2 devices that exhibit signatures of Vth pinning induced by oxide/interface defect states. Finally, we demonstrate FETs using a 2D MoS2 channel and a 6 nm of ZrO2 dielectric, achieving a subthreshold swing of 87 mV dec-1 and a threshold voltage of 0.1 V. Our results offer insights into the role of dielectric/semiconductor interface in 2D MoS2 based FETs for realizing enhancement mode FETs and highlight the potential of ZrO2 as a scalable high-k dielectric.

cond-mat.mtrl-sci

Super amplification of lunar response to gravitational waves driven by thick crust

The Moon has been long regarded as a natural resonator of gravitational waves (GWs) since 1960, showing great potential to fill the frequency gap left behind GW detections by ground- or space-based laser interferometry. However, the spatial variation of this amplification capacity on the Moon remains unclear. Here, we numerically simulate the lunar response to GWs by fully considering the fluctuant topography and laterally heterogeneous interior structures. Our results show that most regions on the Moon can amplify GWs with a ratio over 2, a finding significantly higher than previous estimations. Particularly, the amplification ratio can even reach factors of tens at the resonant frequency of ~0.015 Hz on the highlands surrounding the South Pole-Aitken (SPA) basin, where the regional crust is the thickest. Our findings establish the thick-crust regions as critical zones of GW amplification, which is essential for future landing site selection and instrumental setting for GW detection on the Moon.

astro-ph.EP

LaRe: Latent Refocusing for Multimodal Reasoning

Chain of Thought (CoT) reasoning enhances logical performance by decomposing complex tasks, yet its multimodal extension faces a trade-off. The prevailing Thinking with Images paradigm achieves visual refocusing by explicitly cropping image regions, yet incurs rapidly growing computational overhead. The emerging line of latent-space reasoning reduces token consumption, but lacks the capacity for dynamic refocusing. We argue that this trade-off stems from a tacitly accepted premise that effective visual refocusing must occur in the form of explicit tokens. Building on this, we propose Latent Refocusing (LaRe), a new multimodal reasoning paradigm in which visual refocusing takes place entirely within the latent space. We further design a semantic augmentation training strategy that ensures the semantic structure of the latent space through visual reconstruction objective. Experimental evaluations demonstrate that LaRe improves average accuracy by 7.6% compared to existing baselines while reducing the number of tokens required for inference by 59.7%. When scaled to a 8B-parameter Vision-Language Model backbone, LaRe achieves performance comparable to state-of-the-art methods, demonstrating the efficacy of our proposed latent refocusing paradigm for multimodal reasoning.

cs.CV

Hyperbolic Fracton Model, Subsystem Symmetry and Holography III: Extension to Generic Tessellations

We generalize the Hyperbolic Fracton Model from the $\{5,4\}$ tessellation to generic tessellations, and investigate its core properties: subsystem symmetries, fracton mobility, and holographic correspondence. While the model on the original tessellation has features reminiscent of the flat-space lattice cases, the generalized tessellations exhibit a far richer and more intricate structure. The ground-state degeneracy and subsystem symmetries are generated recursively layer-by-layer, through the inflation rule, but without a simple, uniform pattern. The fracton excitations follow exponential-in-distance and algebraic-in-lattice-size growing patterns when moving outward, and depend sensitively to the tessellation geometry, differing qualitatively from both type-I or type-II fracton model on flat lattices. Despite this increased complexity, the hallmark holographic features -- subregion duality via Rindler reconstruction, the Ryu-Takayanagi formula for mutual information, and effective black hole entropy scaling with horizon area -- remain valid. These results demonstrate that the holographic correspondence in fracton models persists in generic tessellations, and provide a natural platform to explore more intricate subsystem symmetries and fracton physics.

cond-mat.str-el

BachVid: Training-Free Video Generation with Consistent Background and Character

Diffusion Transformers (DiTs) have recently driven significant progress in text-to-video (T2V) generation. However, generating multiple videos with consistent characters and backgrounds remains a significant challenge. Existing methods typically rely on reference images or extensive training, and often only address character consistency, leaving background consistency to image-to-video models. We introduce BachVid, the first training-free method that achieves consistent video generation without needing any reference images. Our approach is based on a systematic analysis of DiT's attention mechanism and intermediate features, revealing its ability to extract foreground masks and identify matching points during the denoising process. Our method leverages this finding by first generating an identity video and caching the intermediate variables, and then inject these cached variables into corresponding positions in newly generated videos, ensuring both foreground and background consistency across multiple videos. Experimental results demonstrate that BachVid achieves robust consistency in generated videos without requiring additional training, offering a novel and efficient solution for consistent video generation without relying on reference images or additional training.

cs.CV

Dual-Space Smoothness for Robust and Balanced LLM Unlearning

As large language models evolve, Machine Unlearning has emerged to address growing concerns around user privacy, copyright infringement, and overall safety. Yet state-of-the-art (SOTA) unlearning methods often suffer from catastrophic forgetting and metric imbalance, for example, by over-optimizing one objective (e.g., unlearning effectiveness, utility preservation, or privacy protection) at the expense of others. In addition, small perturbations in the representation or parameter space can be exploited by relearn and jailbreak attacks. To address these challenges, we propose PRISM, a unified framework that enforces dual-space smoothness in representation and parameter spaces to improve robustness and balance unlearning metrics. PRISM consists of two smoothness optimization stages: (i) a representation space stage that employs a robustly trained probe to defend against jailbreak attacks, and (ii) a parameter-space stage that decouples retain-forget gradient conflicts, reduces imbalance, and smooths the parameter space to mitigate relearning attacks. Extensive experiments on WMDP and MUSE, across conversational-dialogue and continuous-text settings, show that PRISM outperforms SOTA baselines under multiple attacks while achieving a better balance among key metrics.

cs.CL