Search arXivSearch

arXiv subjects

Yaping Li

Publications and source records attributed to Yaping Li.

12 recordsLinked to original sources

KAI: A Kinematic-Aware Interface for Data-Efficient Articulated Object Manipulation

Articulated object manipulation requires an understanding of kinematic structure that is difficult and costly to learn from robot demonstrations alone. We introduce the Kinematic-Aware Articulation Interface (KAI), a structured intermediate representation that captures the kinematic structure of articulated objects. By embedding interpretable geometric and kinematic priors into policy learning, KAI provides a strong inductive bias aligned with the underlying structure of articulated motion. This design effectively improves sample efficiency, with gains particularly pronounced in low-data regimes: across six simulation tasks, our method achieves an average success rate of 82.9%, matching or surpassing baseline performance while using only half the demonstration data. Our method also exhibits robust generalization to unseen backgrounds and visual distractors, transferring from a single clean training environment to cluttered real-world scenes. KAI's action-agnostic design further enables co-training with human interaction videos to enhance real-world robustness: under diverse visual distractions, our method with video co-training achieves over 70% average success rate.

cs.RO

Reasoning for Mobile User Experience with Multimodal LLMs: Task, Benchmark, and Approach

User experience (UX) centered on usability, perceived consistency, and functional clarity is fundamental to real-world user interfaces (UI). The application of multimodal large language models (MLLMs) in the field of user interfaces is evolving rapidly, such as visual element grounding, graphical user interface (GUI) agents, and design-to-code generation. However, research efforts on evaluating UX based on UI screenshots are still immature. To address this, we propose UXBench, a novel multimodal benchmark consisting of 2,000 VQA data samples designed to assess MLLMs' ability to perform UI-based reasoning. UXBench includes 8 tasks based on real-world UI screenshots that require fine-grained diagnosis of UX issues across layout relationships, visual hierarchy, and content consistency. Our extensive evaluation of mainstream MLLMs shows that they remain fundamentally limited in their capacity for UI-based reasoning. The results underscore the need for further advancements in this area. To bridge this gap, we propose UI-UX, an MLLM based on Qwen3-VL-4B-Thinking foundation model and enhanced via reinforcement learning with two key innovations: a reward routing mechanism that dynamically balances perceptual understanding and logical reasoning during inference, and an asymmetric transition reward that suppresses redundant or insufficient reasoning steps. Experiments demonstrate that UI-UX achieves state-of-the-art (SOTA) performance on UXBench, attaining an accuracy of 0.7963 -- surpassing Claude-4.5-Sonnet's 0.6550 -- while exhibiting strong generalization across diverse UI tasks and maintaining low inference latency.

cs.AI

DASH: Fast Differentiable Architecture Search for Hybrid Attention in Minutes on a Single GPU

Hybrid attention architectures are becoming an increasingly important paradigm for improving LLM inference efficiency while preserving model quality, making hybrid architecture design a central problem. Existing designs often rely on manual empirical rules or proxy-based selector signals for layer-wise operator allocation. Recent NAS-style systems such as Jet-Nemotron demonstrate the promise of automated hybrid architecture search. However, Jet-Nemotron's PostNAS search stages alone use 200B tokens, making such search pipelines difficult to use as routine methods for hybrid architecture design. We introduce DASH, a fast differentiable search framework for hybrid attention architecture design, which relaxes discrete layer-wise attention operator placement into continuous architecture logits, prepares reusable teacher-aligned linear candidates, and performs architecture-only search with model and operator weights frozen to significantly enhance search efficiency. On Qwen2.5-3B-Instruct, DASH consistently outperforms a comprehensive suite of existing selector-style hybrid attention design baselines, showing that direct differentiable search can discover stronger hybrid architectures. Moreover, DASH achieves stronger RULER performance than released Jet-Nemotron models while remaining competitive on overlapping short-context and general benchmarks. Notably, each DASH search run uses only 12.3M tokens and takes about 20 minutes on a single RTX Pro 6000 GPU, corresponding to merely 0.006% of the PostNAS search tokens reported by Jet-Nemotron. These results suggest that high-quality hybrid attention architectures can be obtained through minutes-level differentiable search, providing a promising direction for hybrid architecture design.

cs.LG

Residual-Squeezing Mechanism of Mismatch in Inverse-Squeezing Kennedy Receivers

The discrimination of quantum states is fundamental to quantum information processing. Inverse-squeezing Kennedy (IS-Kennedy) receivers can outperform the coherent-state BPSK Helstrom benchmark at the same energy by converting transmitter-side squeezing into an effective coherent-state separation gain, without violating the Helstrom bound for the squeezed-state alphabet. This work investigates how squeezing mismatch degrades this mechanism. We show that imperfect inverse squeezing transforms the ideally nulled output into a residually squeezed state, thereby altering the photon-number statistics before detection. This residual-squeezing picture reveals a strong physical asymmetry between squeezing-magnitude and squeezing-phase mismatches. Magnitude mismatch produces an energy-independent error floor in the high-signal-energy regime, whereas phase mismatch generates a residual squeezing term that grows with signal energy. In the small-residual-squeezing regime, this leads to a polynomial growth of the leading error contribution and a rapid collapse of the SQL advantage. We also identify a parity-step effect in photon-number-resolving detection: because the nulled residual squeezed vacuum contains only even photon numbers, increasing detector resolution improves the high-energy robustness only when the effective saturation threshold crosses the next even photon number. These results identify phase locking as the dominant bottleneck for IS-Kennedy-type non-Gaussian receivers under unitary squeezing mismatch and provide design guidelines for robust squeezed-state quantum receivers.

quant-ph

Near-optimal discrimination of displaced squeezed binary signals using displacement, inverse-squeezing, and photon-number-resolving detection

We propose an inverse-squeezing Kennedy receiver for discriminating binary phase-shift-keyed displaced squeezed vacuum states. The receiver combines a Kennedy-type nulling displacement, an orthogonally oriented inverse-squeezing operation and photon-number-resolving detection with a maximum-a-\emph{posteriori} threshold rule. Its key mechanism is that the inverse-squeezing stage converts transmitter-side squeezing into enhanced photon-number contrast, or equivalently an effective coherent-state energy gain, that can be directly exploited at the measurement stage. Under ideal equal-prior conditions, the receiver surpasses the standard quantum limit for squeezed-state binary phase-shift keying at approximately $N\approx 0.3$, outperforms the Helstrom bound of coherent-state binary phase-shift keying at approximately $N\approx 0.4$, and reaches the 1\% error level near $N\approx 0.6$. We further analyze its performance under realistic imperfections, including finite detector efficiency, dark counts, channel phase diffusion, receiver thermal noise and transmission loss. The results show that adaptive thresholding preserves robust performance against detector and noise imperfections over practical parameter ranges, whereas transmission loss progressively suppresses the squeezing-enabled advantage. These findings indicate that, for the fixed source parametrization adopted in this work, the proposed receiver is most advantageous in the low-loss regime, especially at low source energies.

quant-ph

InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy

Recent works explore how real and synthetic data contribute to Vision-Language-Action (VLA) models' generalization. While current VLA models have shown the strong effectiveness of large-scale real-robot pre-training, synthetic data has not previously demonstrated comparable capability at scale. This paper provides the first evidence that synthetic data alone can match the performance of the strongest $\pi$-dataset in pre-training a VLA model, revealing the substantial value of large-scale simulation. The resulting model also exhibits surprisingly zero-shot sim-to-real transfer on several challenging tasks. Our synthetic dataset, InternData-A1, contains over 630k trajectories and 7,433 hours across 4 embodiments, 18 skills, 70 tasks, and 227 scenes, covering rigid, articulated, deformable, and fluid-object manipulation. It is generated through a highly autonomous, fully decoupled, and compositional simulation pipeline that enables long-horizon skill composition, flexible task assembly, and heterogeneous embodiments with minimal manual tuning. Using the same architecture as $\pi_0$, we pre-train a model entirely on InternData-A1 and find that it matches the official $\pi_0$ across 49 simulation tasks, 5 real-world tasks, and 4 long-horizon dexterous tasks. We release the dataset and will open-source the generation pipeline to broaden access to large-scale robotic data and to lower the barrier to scalable data creation for embodied AI research.

cs.RO

Oxygen vacancies modulated VO2 for neurons and Spiking Neural Network construction

Artificial neuronal devices are the basic building blocks for neuromorphic computing systems, which have been motivated by realistic brain emulation. Aiming for these applications, various device concepts have been proposed to mimic the neuronal dynamics and functions. While till now, the artificial neuron devices with high efficiency, high stability and low power consumption are still far from practical application. Due to the special insulator-metal phase transition, Vanadium Dioxide (VO2) has been considered as an idea candidate for neuronal device fabrication. However, its intrinsic insulating state requires the VO2 neuronal device to be driven under large bias voltage, resulting in high power consumption and low frequency. Thus in the current study, we have addressed this challenge by preparing oxygen vacancies modulated VO2 film(VO2-x) and fabricating the VO2-x neuronal devices for Spiking Neural Networks (SNNs) construction. Results indicate the neuron devices can be operated under lower voltage with improved processing speed. The proposed VO2-x based back-propagation SNNs (BP-SNNs) system, trained with the MNIST dataset, demonstrates excellent accuracy in image recognition. Our study not only demonstrates the VO2-x based neurons and SNN system for practical application, but also offers an effective way to optimize the future neuromorphic computing systems by defect engineering strategy.

cs.NE

STURE: Spatial-Temporal Mutual Representation Learning for Robust Data Association in Online Multi-Object Tracking

Online multi-object tracking (MOT) is a longstanding task for computer vision and intelligent vehicle platform. At present, the main paradigm is tracking-by-detection, and the main difficulty of this paradigm is how to associate current candidate detections with historical tracklets. However, in the MOT scenarios, each historical tracklet is composed of an object sequence, while each candidate detection is just a flat image, which lacks temporal features of the object sequence. The feature difference between current candidate detections and historical tracklets makes the object association much harder. Therefore, we propose a Spatial-Temporal Mutual Representation Learning (STURE) approach which learns spatial-temporal representations between current candidate detections and historical sequences in a mutual representation space. For historical trackelets, the detection learning network is forced to match the representations of sequence learning network in a mutual representation space. The proposed approach is capable of extracting more distinguishing detection and sequence representations by using various designed losses in object association. As a result, spatial-temporal feature is learned mutually to reinforce the current detection features, and the feature difference can be relieved. To prove the robustness of the STURE, it is applied to the public MOT challenge benchmarks and performs well compared with various state-of-the-art online MOT trackers based on identity-preserving metrics.

cs.CV

Ab initio calculations of the atomic and electronic structures of crystalline PEO3:LiCF3SO3 electrolytes

The PEO3:LiCF3SO3 polymer electrolyte has attracted significant research due to its enhanced stability at the lithium/polymer interface of high conductivity polymer batteries. Experimental studies have shown that, depending on the preparation conditions, both the PEO3:LiCF3SO3 crystalline complex and the PEO3:LiCF3SO3 amorphous phase can be formed. However, previous theoretical investigations focused on the short chain amorphous PEO3:LiCF3SO3 system. We report ab initio density-functional-theory calculations of crystalline PEO3:LiCF3SO3. The calculated results about the bonding configuration, electronic structures, and conductivity properties are in good agreement with the experimental measurements.

cond-mat.mtrl-sci

Evolution Process of Wurtzite ZnO Films on Cubic MgO (001) Substrates: a Structural, Optical and Electronic Investigation of the Misfit Structures

The interface between hexagonal ZnO films and cubic MgO (001) substrates, fabricated through molecular beam epitaxy, are thoroughly investigated. X-ray diffraction and (scanning) transmission electron microscopy reveal that, at the substrate temperature above 200 degree C, the growth follows the single [0001] direction; while at the substrate below 150 degree C, the growth is initially along [0001] and then mainly changes to [0-332] variants beyond the thickness of about 10 nm. Interestingly, a double-domain feature with a rotational angle of 30 degree appears for the growth along [0001] regardless of the growth temperature, experimentally demonstrated the theoretical predictions for occurrence of double rotational domains in such a heteroepitaxy [Grundmann et al, Phys. Rev. Lett. 105, 146102 (2010)]. It is also found that, the optical transmissivity of the ZnO film is greatly influenced by the mutation of growth directions, stimulated by the bond-length modulations, as further determined by X-ray absorption Spectra (XAS) at Zn K edge. The XAS results also show the evolution of 4pxy and 4pz states in the conduction band as the growth temperature increases. The results obtained from this work can hopefully promote the applications of ZnO in advanced optoelectronics for which its integration with other materials of different phases is desirable.

cond-mat.mtrl-sci

The non-ballistic superluminal motion in the plane of the sky-II

The model of non-ballistic jet motion proposed in 2008 provides a simple explanation to the inward jet motion and bent jet. Recently, evidences of such a non-radial motion increase rapidly, and more complicated morphologies appear. On the other hand, the ballistic plus precession model likely holds in majority samples of jet motion. This paper discusses the relationship between the ballistic and non-ballistic model of jet motion, which suggests that the interaction of ejectors with ambient matter can produce knots at different stages of evolution and hence different separations to the core. And as a jet precesses, knots produced between the core and the deceleration radius result in spiral pattern expected by the model of ballistic plus precession; and knots generated at the deceleration radius display non-radial motion such as bent jet or oscillation of ridge-line. This paper develops the first non-ballistic model in four aspects. Firstly, it provides a numerical simulation to the production of multi-knot for a precessing jet. Secondly, it fits the precession behavior of multi-knot and interprets the oscillation of ridge lines like S5 1803+784. Thirdly, it gives an unified interpretation to the bent jet applicable to both multi-knot and single knot. And fourthly, the problem of very large numbers of observed outward motions as opposed to the inward ones is addressed in a new scope.

astro-ph.HE

Enabling Multi-level Trust in Privacy Preserving Data Mining

Privacy Preserving Data Mining (PPDM) addresses the problem of developing accurate models about aggregated data without access to precise information in individual data record. A widely studied \emph{perturbation-based PPDM} approach introduces random perturbation to individual values to preserve privacy before data is published. Previous solutions of this approach are limited in their tacit assumption of single-level trust on data miners. In this work, we relax this assumption and expand the scope of perturbation-based PPDM to Multi-Level Trust (MLT-PPDM). In our setting, the more trusted a data miner is, the less perturbed copy of the data it can access. Under this setting, a malicious data miner may have access to differently perturbed copies of the same data through various means, and may combine these diverse copies to jointly infer additional information about the original data that the data owner does not intend to release. Preventing such \emph{diversity attacks} is the key challenge of providing MLT-PPDM services. We address this challenge by properly correlating perturbation across copies at different trust levels. We prove that our solution is robust against diversity attacks with respect to our privacy goal. That is, for data miners who have access to an arbitrary collection of the perturbed copies, our solution prevent them from jointly reconstructing the original data more accurately than the best effort using any individual copy in the collection. Our solution allows a data owner to generate perturbed copies of its data for arbitrary trust levels on-demand. This feature offers data owners maximum flexibility.

cs.DB