Search arXivSearch

arXiv subjects

Xuan Han

Publications and source records attributed to Xuan Han.

8 recordsLinked to original sources

DynamicWAM: Dual-Path Motion Conditioning for World-Action Models in Dynamic Manipulation

Dynamic manipulation requires robots to infer target motion and respond promptly, yet existing World-Action Models (WAMs) typically condition only on the current frame and execute large backbones synchronously, limiting motion awareness and responsive control in dynamic scenes. We propose DynamicWAM, a compact WAM for dynamic object manipulation with dual-path motion conditioning. DynamicWAM introduces history-flow conditioning, encoding temporally aligned optical-flow frames alongside the current observation through a frozen pretrained video VAE to preserve spatial motion structure, while injecting kinematic descriptors of displacement, duration, velocity, and acceleration into the action expert to provide motion magnitude and timing. The two complementary paths are fused through joint world-action attention. A distilled compact backbone and real-time chunking (RTC)-based asynchronous execution further enable responsive control. On DOMINO, DynamicWAM achieves a 38.2% success rate and a 53.2 manipulation score, outperforming all evaluated baselines. Across 12 real-world tasks spanning linear, circular, and compound target motion, it achieves a 46.7% average success rate, exceeding the strongest baseline by 22.9 percentage points.

cs.RO

Improving Text-Instance Alignment Of Foreground Conditioned Out-Painting Via Customized Concept Embedding

To showcase products, merchants often incur substantial costs creating high-quality display images. Foreground Conditioned Outpainting (FCO) meets this demand, allowing users to create desired backgrounds for foreground instances at a low cost by adjusting the text prompt. However, existing text-driven FCO methods exhibit critical flaws in their outputs, most notably the presence of artifacts, which refer to regions in the synthesized background that share the same semantics as the foreground instance. Such artifacts diminish the object's prominence and degrade image quality. We attribute the issue to the misalignment between the given instance and text-derived concept embeddings. To address this, we propose the Customized Concept Embedding Diffusion (CCE-Diffusion) framework. Its core is a CCE-Module to customize concept embeddings, bridging the gap between generic noun semantics and a specific visual instance. An Instance-Aware Loss guides the module's optimization, while a Semantic-Preserving Prompt Template prevents customized embeddings from distorting other words in the prompt. Both qualitative and quantitative evaluations demonstrate that CCE-Diffusion significantly reduces artifacts in the outputs. As a plug-and-play component, the CCE-Module can integrate with various FCO methods, enhancing their performance.

cs.CV

Pose-ICL: 3D-Aware In-Context Learning for Pose-Controllable Subject Customization

Subject Customization is a foundational task in modern image generation. By providing a few reference images and a text prompt, users can generate images of a specific object in any desired scene. However, existing methods still struggle to achieve effective pose control for customized subjects. In practice, they often exhibit inaccurate poses or inconsistent cross-pose appearances. These limitations suggest that understanding objects in a volumetric manner remains a significant challenge for 2D-native backbones. To address this challenge, we propose Pose-ICL, a tuning-free framework that leverages 3D-aware In-Context Learning (ICL) to directly adapt to new subjects through multiple paired image-pose references. Its core mechanism,Surface-Anchored Position Embedding (SAPE), equips the model with explicit 3D awareness by anchoring image tokens to the surface coordinates of a volumetric bounding box. Dedicated refinements ensure its seamless compatibility with existing DiT models. Extensive evaluations on both 3D assets and real-world subjects demonstrate that Pose-ICL significantly outperforms current methods in both pose accuracy and identity consistency.

cs.CV

3D Part Assembly Generation with Instance Encoded Transformer

It is desirable to enable robots capable of automatic assembly. Structural understanding of object parts plays a crucial role in this task yet remains relatively unexplored. In this paper, we focus on the setting of furniture assembly from a complete set of part geometries, which is essentially a 6-DoF part pose estimation problem. We propose a multi-layer transformer-based framework that involves geometric and relational reasoning between parts to update the part poses iteratively. We carefully design a unique instance encoding to solve the ambiguity between geometrically-similar parts so that all parts can be distinguished. In addition to assembling from scratch, we extend our framework to a new task called in-process part assembly. Analogous to furniture maintenance, it requires robots to continue with unfinished products and assemble the remaining parts into appropriate positions. Our method achieves far more than 10% improvements over the current state-of-the-art in multiple metrics on the public PartNet dataset. Extensive experiments and quantitative comparisons demonstrate the effectiveness of the proposed framework.

cs.RO

Quantum Random Number Generation with Uncharacterized Laser and Sunlight

The entropy or randomness source is an essential ingredient in random number generation. Quantum random number generators generally require well modeled and calibrated light sources, such as a laser, to generate randomness. With uncharacterized light sources, such as sunlight or an uncharacterized laser, genuine randomness is practically hard to be quantified or extracted owing to its unknown or complicated structure. By exploiting a recently proposed source-independent randomness generation protocol, we theoretically modify it by considering practical issues and experimentally realize the modified scheme with an uncharacterized laser and a sunlight source. The extracted randomness is guaranteed to be secure independent of its source and the randomness generation speed reaches 1 Mbps, three orders of magnitude higher than the original realization. Our result signifies the power of quantum technology in randomness generation and paves the way to high-speed semi-self-testing quantum random number generators with practical light sources.

quant-ph

Point-ahead demonstration of a transmitting antenna for satellite quantum communication

A low-divergence beam is an essential prerequisite for a high-efficiency longdistance optical link, particularly for satellite-based quantum communication. A point-ahead angle, caused by satellite motion, is always several times larger than the divergence angle of the signal beam. We design a novel transmitting antenna with a point-ahead function, and provide an easy-to-perform calibration method with an accuracy better than 0.2 urad. Subsequently, our antenna establishes an uplink to the quantum satellite, Micius, with a link loss of 41-52 dB over a distance of 500-1,400 km. The results clearly confirm the validity of our model, and provide the ability to conduct quantum communications. Our approach can be adopted in various free space optical communication systems between moving platforms.

quant-ph

Polarization design for ground-to-satellite quantum entanglement distribution

Polarization maintenance is a key technology for free-space quantum communication. In this paper, we describe a polarization maintenance design of a transmitting antenna with an average polarization extinction ratio of 887 : 1 by a local test. We implemented a feasible polarization-compensation scheme for satellite motions that has a polarization fidelity more than 0.995. Finally, we distribute entanglement to a satellite from ground for the first time with a violation of Bell inequality by 2.312+-0.096.

quant-ph

Ground-to-satellite quantum teleportation

An arbitrary unknown quantum state cannot be precisely measured or perfectly replicated. However, quantum teleportation allows faithful transfer of unknown quantum states from one object to another over long distance, without physical travelling of the object itself. Long-distance teleportation has been recognized as a fundamental element in protocols such as large-scale quantum networks and distributed quantum computation. However, the previous teleportation experiments between distant locations were limited to a distance on the order of 100 kilometers, due to photon loss in optical fibres or terrestrial free-space channels. An outstanding open challenge for a global-scale "quantum internet" is to significantly extend the range for teleportation. A promising solution to this problem is exploiting satellite platform and space-based link, which can conveniently connect two remote points on the Earth with greatly reduced channel loss because most of the photons' propagation path is in empty space. Here, we report the first quantum teleportation of independent single-photon qubits from a ground observatory to a low Earth orbit satellite - through an up-link channel - with a distance up to 1400 km. To optimize the link efficiency and overcome the atmospheric turbulence in the up-link, a series of techniques are developed, including a compact ultra-bright source of multi-photon entanglement, narrow beam divergence, high-bandwidth and high-accuracy acquiring, pointing, and tracking (APT). We demonstrate successful quantum teleportation for six input states in mutually unbiased bases with an average fidelity of 0.80+/-0.01, well above the classical limit. This work establishes the first ground-to-satellite up-link for faithful and ultra-long-distance quantum teleportation, an essential step toward global-scale quantum internet.

quant-ph