Search arXivSearch

arXiv subjects

Jing Wu

Publications and source records attributed to Jing Wu.

At least 19 recordsLinked to original sources

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whose trainable generative and action components are all initialized from scratch on manipulation data. It combines a control-oriented autoencoder (CoAE), a single-step visual planner (SVP), and an inverse dynamics model (IDM). CoAE retains action- and instruction-relevant information under aggressive compression, while SVP produces a complete future state in one differentiable pass, so visual planning and inverse dynamics can be pretrained separately on complementary data. The components are then jointly trained with knowledge-aligned selective optimization (KASO), which reduces mismatched supervision by selecting only predicted futures judged behaviorally compatible with the recorded action. We evaluate pretrained checkpoints directly, without per-task fine-tuning, on 100 tasks across 20 manipulation skill groups with held-out scenes, backgrounds, lighting, and object instances. Scaling co-training data from 300 to 30,000 hours raises success from 17.1% to 44.1% on G1-OP and from 13.4% to 31.1% on G2-90D; despite comprising less than 2% of the co-training data, G2-90D improves by 17.7 points, suggesting cross-embodiment transfer. Gains span 19/20 and 18/20 skill groups, and skill-specific coverage strongly correlates with zero-shot out-of-distribution (OOD) success (Pearson r=0.80; Spearman rho=0.85). Under the same protocol, the model grounds object, color, shape, and position references in at least 90% of trials and follows explicit instructions even when they conflict with an already-committed behavior or a conventional scene association.

cs.RO

Magnetic Moments and Radiative Transitions of the $T_{cc}(3875)^+$ and its partner states

We systematically investigate the magnetic moments (MMs) and radiative decay widths of S-wave doubly heavy tetraquark states based on chromomagnetic interaction (CMI). They provide a complementary probe to distinguish compact from molecule structures in addition to the mass spectrum and strong decay properties. Using the CMI eigenvectors, we compute the MMs and M1 transition rates for the tetraquarks in the compact configuration. We predict the MM of the observed $T_{cc}(3875)^+$ with $I(J^P)=0(1^+)$ to be 0.45 $\mu_N$ when treating it as a compact tetraquark state, whereas the MM of the $0(1^+)$ $DD^*$ molecule is about $-0.07\,\mu_N$. Five radiative transition channels related to the $T_{cc}(3875)^+$ are identified with widths ranging from $6.07$ keV to $306.37$ keV. Our results show that the MMs of the $J^P=1^+$ $bc\bar{q}\bar{q}^\prime$ ($q/q^\prime=u,d,s$) and $J^P=1^+$ $QQ\bar{n}\bar{s}$ ($Q=b,c; n=u,d$) states are influenced by the diquark-spin mixing. Through the analyses of radiative transitions between different tetraquark states, we find that such processes in the $QQ\bar{n}\bar{n}^\prime$, $QQ\bar{s}\bar{s}$, and $bc\bar{n}\bar{s}$ cases may serve to reveal the tetraquark structures of the initial or final states. We also define the magnetic coupling matrices characterizing the MMs of the tetraquark system with $J^P=1^+$, with which the range of MM can be constrained. The present study provides a valuable reference point for the search of exotic states in future particle physics experiments.

hep-ph

SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or external spatial encoders, which increase complexity and degrade the underlying VLMs' general-purpose capabilities after spatial fine-tuning. To this end, we propose a parameter-efficient \textit{\textbf{Spatio}-vision \textbf{L}anguage \textbf{M}odels (SpatioLM)}, that enhances spatial intelligence without extra 3D prior inputs or third-party spatial encoders. Concretely, we design a plug-and-play and non-invasive spatio-vision module that elicits the spatial knowledge inherent in VLMs. Furthermore, we innovatively leverage pseudo depth and camera information as supervision to guide the model in learning physically coherent representations. Extensive experiments show that SpatioLM achieves significant improvements in diverse tasks, including spatial perception and understanding while effectively limiting the degradation of general capabilities. Notably, the model achieves an impressive score of 71.6 on the VSI-Bench (the first model to surpass 70). In addition, it attains competitive performance when transferred to embodied manipulation tasks. Code is available at \href{https://github.com/xiaomi-research/spatio-lm}{\faGithub~spatio-lm}.

cs.CV

Object Detection for Autonomous Driving in Chinese Rural Scenes: An Experimental Study on Real-Synthetic Data Mixing and Model Evaluation

Currently, autonomous driving object detection models face significant data scarcity and generalization challenges when navigating complex Chinese rural traffic scenarios. To address these limitations, we propose a novel real-synthetic mixed object detection dataset tailored specifically for Chinese rural roads and systematically evaluate the performance of 13 mainstream detectors under different real-to-synthetic data ratios, thereby providing empirical evidence for model selection and data strategy design in rural autonomous driving scenarios. Our dataset combines real-world images captured in Weishi County, Henan Province, with parameterized virtual scenes generated via Unreal Engine. To accurately reflect the unique realities of rural traffic, we define a comprehensive 14-category object system encompassing region-specific elements such as electric tricycles, low-speed vehicles (LSVs), and roadside stalls. Under a unified training protocol, we systematically evaluate 13 mainstream detectors -- spanning the YOLOv5, YOLOv8, YOLO11, and YOLO26 series, as well as RT-DETR-L -- across three data configurations: an all-real baseline, a 1:0.5 real-to-virtual mix, and a 1:1 mix. Experimental results demonstrate that a moderate injection of synthetic data (1:0.5 ratio) effectively enhances detection performance, with YOLO11m achieving the highest mAP@0.5 of 0.758. However, a higher proportion of synthetic data (1:1) introduces domain shifts that offset the benefits of data scaling. While most models reliably identify distinct local vehicles, significant perceptual bottlenecks remain for long-tail, non-standard objects like stalls and railings. This research provides crucial empirical evidence and novel insights for model selection and synthetic data strategies, facilitating the practical deployment of autonomous driving perception systems in rural areas.

cs.CV

Universal Optimization and Tighter Fidelity Bounds for Approximate Quantum Error Correction

Approximate quantum error correction (AQEC) extends the framework of discrete- and continuous-variable quantum error correction beyond the Knill-Laflamme (KL) conditions, where the recovery performance is quantified by entanglement fidelity. Recent studies have enabled efficient evaluation of near-optimal entanglement fidelity using transpose-channel recovery. Yet, determining the global optimal recovery map and its entanglement fidelity for general codes beyond the KL conditions remains a major computational challenge. Direct optimization becomes prohibitive as the number of noise Kraus operators grows rapidly with system size, and existing approaches lack rigorous guarantees for reducing this optimization to a tractable dimension. Here, we derive an explicit characterization of the optimal environmental state of complement channel, which transforms the optimization over recovery channels into an equivalent optimization over quotient unitaries. For a broader class of codes that satisfy only the orthogonality part of the KL conditions, we show that the optimal recovery map admits an explicit analytical form. Building on this form, we derive novel rigorous lower bounds of entanglement fidelity that strictly improve upon the transpose-recovery bound. We further develop a novel recovery strategy based on principle components, and derive a rigorous bound on the error introduced by noise truncation. Our approach enables efficient searches for approximate recovery maps for AQEC codes, avoiding the need to optimize over the full Kraus-operator space.

quant-ph

Quantitative Homogenization of PDEs with Neumann boundary conditions: a probabilistic approach

In this paper, we study quantitative homogenization for viscosity solutions of multi-scale semilinear second order partial differential equations (PDEs) on convex domains with Neumann boundary conditions. To this aim we use the probabilistic approach by studying the quantitative homogenization of backward stochastic differential equations (SDEs) associated with slow-fast systems of reflected SDEs.

math.PR

Negative temperature coefficient of Gilbert damping in magnetic bilayers

The Gilbert damping of magnetic materials is an important magnetic parameter that determines the switching speed and energy dissipation of spintronic devices. In simple metals, the intrinsic Gilbert damping increases with temperature and diverges near the Curie temperature as a result of spin fluctuations. Here we present atomistic simulations and experimental measurements showing surprising and opposite behavior in Py/Nd bilayers, where the Gilbert damping decreases with increasing temperature. The effect arises because of the enhanced damping at the interface as a result of spin pumping, where elevated temperatures cause a dynamic separation of the interfacial and bulk magnetization during relaxation. Furthermore, the temperature dependence of the damping can be controlled by varying the thickness of the Nd capping layer. Our findings present a new spintronic effect that can be used to modify the dynamic properties of nanoscale materials and devices for enhanced energy efficiency or with improved switching dynamics.

cond-mat.mes-hall

Functional renormalization group study of the jet quenching parameter near the QCD critical end point

We investigate the jet quenching parameter $\hat{q}$ in the QCD phase diagram within a QCD-assisted low-energy effective theory using the functional renormalization group (fRG). Following the formalism that relates $\hat{q}$ to the spectral functions of the chiral order-parameter field, we compute the $\sigma$ and $\pi$ meson contributions to $\hat{q}$ at finite temperature and baryon chemical potential from analytically continued mesonic two-point functions. We find that $\hat{q}$ receives appreciable contributions mainly above the chiral phase boundary and exhibits a pronounced enhancement at large baryon chemical potential as the chiral crossover sharpens toward the critical end point (CEP), a behavior consistent with the picture of partonic critical opalescence (PCO), a pronounced enhancement of jet transverse momentum broadening induced by the critical $\sigma$ field fluctuations.

hep-ph

SurgRFO: Foundation Model Based Compositional Synthesis of Critical Retained Foreign Objects in Intraoperative Chest X-rays

Critical retained foreign objects (RFOs) on intraoperative chest radiographs are rare but high-risk events. Their scarcity limits robust automated detection model training and generalization. We introduce SurgRFO, a two-stage synthesis framework for generating realistic RFO-present intraoperative chest X-rays. In Stage 1, a Roentgen chest X-ray foundation model is fine-tuned on surgical-domain images to generate realistic RFO-free backgrounds that preserve anatomy, indwelling lines and tubes, and intraoperative imaging characteristics. In Stage 2, a lightweight generator trained on localized RFO patches from limited positive cases synthesizes diverse RFO instances, which are composited onto generated backgrounds using conditional Poisson fusion to improve photometric consistency. We evaluate SurgRFO through (i) a blinded clinician study assessing realism and clinical plausibility, and (ii) downstream detection experiments in which synthesized data are used to augment Faster R-CNN, YOLOv8, and RetinaNet. SurgRFO consistently improves sensitivity at low false-positive-per-image (FPPI) operating points on internal and external test sets. Clinician ratings indicate that the synthesized images achieve realism comparable to real intraoperative images. Ablation analyses further examine fusion strategies and synthesis scale. Ethical safeguards for synthetic surgical data are also discussed.

eess.IV

Self-Supervised Contrastive Learning for Cardiac MR Sequence Classification

Vision Transformer (ViT) models, utilizing self-attention mechanisms, have demonstrated robust generalization capabilities across various vision tasks, including image classification. However, these models, typically pretrained on general public datasets, often lack the specialized domain knowledge necessary for medical imaging applications. In this study, we investigate the adaptation of ViT models, specifically for cardiac magnetic resonance (MR) images, using an in-house dataset. We found that pretrained ViT features do not effectively transfer to the cardiac MR domain. To overcome this limitation, we introduce an adaptation strategy that utilizes image-based self-supervised contrastive learning, demonstrating superior performance compared to traditional supervised training approaches. Moreover, our adapted ViT model exhibits strong generalization to external MR datasets such as BraTS and ADNI. Through ablation studies, we further investigate the impact of batch size and dataset scale on performance. Ultimately, our adapted model achieves classification AUC exceeding 0.75 across the four most common cardiac MR sequences.

cs.CV

AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture

Agricultural decision-making increasingly requires multimodal systems that can transform visual observations into reliable, executable actions. However, existing agricultural multimodal benchmarks mainly evaluate final-answer correctness and provide limited support for assessing whether models can use external tools to complete precision-sensitive workflows. In this paper, we introduce AgroTools, a benchmark for evaluating tool-augmented multimodal agents in agriculture. AgroTools contains 539 question-answer instances paired with 1,097 heterogeneous agricultural images, spanning five task families and an executable environment of 14 agricultural tools. Each query is annotated with structured tool-use traces, enabling a dual-view evaluation of both process-level execution quality and outcome-level task success. We benchmark 9 open-source and 4 closed-source multimodal large language models on AgroTools. Results show that current models remain far from reliable in agricultural tool-use settings, with clear bottlenecks in tool planning, argument generation, execution recovery, and final-answer synthesis. We hope AgroTools will support future research on multimodal agents for high-precision agricultural applications. The benchmark and evaluation are available at https://huggingface.co/datasets/AgroTools/AgroTools.

cs.CV

Antiferromagnetic Ordering Enhanced Magnetic Damping in Mn2Au/CoFeB Bilayers

Antiferromagnets (AFMs) hold significant potential for spintronic devices owing to their insensitivity to external magnetic fields and the absence of stray fields. Beyond these inherent advantages, an AFM can manipulate the magnetic dynamics of a ferromagnet (FM) layer in AFM/FM bilayers, whereas the mechanism of such manipulation remains controversial. Here, we investigate the magnetic dynamics of AFM/FM Mn2Au/CoFeB bilayers via Ferromagnetic Resonance (FMR). It is found that the N\'eel temperature of 2-nm-thick Mn2Au is as low as ~40 K, in sharp contrast to that of bulk Mn2Au, which exceeds 1000 K. In the Mn2Au(2 nm)/CoFeB(4 nm) bilayer, the magnetic damping ${\alpha}$ of the CoFeB layer increases from 0.013 to 0.047 as temperature decreases from 160 K to 10 K, accompanied by a synchronous increase in the exchange coupling field H_rot. Such an increase in ${\alpha}$ is attributed to the enhanced spin angular momentum transfer from CoFeB to Mn2Au, mediated through AFM-FM exchange coupling between Mn2Au and CoFeB, which is enhanced by the Mn2Au antiferromagnetic ordering as the temperature decreases. Our study provides deeper insights into AFM/FM dynamics and spintronic storage technology.

cond-mat.mtrl-sci

Engineering Hybrid Resonances in Nanophotonics

Hybridization of resonances is known to overcome inherent limitations of individual systems, enabling advanced functionalities and applications. Here we discuss hybrid plasmonic-Mie resonators that emerged recently as a promising direction in advancing nanophotonic structures by synergistically combining the strong near-field enhancement of plasmonic components with the low-loss, multipolar resonances of dielectric Mie elements. We review the recent progress in the field, encompassing the fundamental physical principles, structural design strategies, material platforms, computational optimization approaches, and representative device implementations. Our discussion starts by evaluating the complementary characteristics of plasmonic and Mie resonances followed by a description of the coupling between these resonances in order to boost light-matter interactions. Afterward, we explore the performance of efficient hybrid resonators for different application areas. Apart from the conventional metal-dielectric systems, we consider the recent class of epsilon-near-zero (ENZ) materials, which can provide unique advantages in terms of field localization, phase engineering, and energy flow management in the vicinity of zero-permittivity conditions, offering more flexibility in designing hybrid nano-optical devices. Lastly, we point out potential research avenues aiming to improve functional and efficient nanophotonic devices, especially those involving emerging topological material systems, such as Sb2Te3, Bi2Te3, Bi2Se3, combining plasmonic amplification, dielectric confinement, and spin-dependent optical behavior.

physics.optics

JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR

Reinforcement learning with verifiable rewards (RLVR) enhances the reasoning of large language models (LLMs), but standard RLVR often depends on human-annotated answers or carefully curated reward specifications. In machine-checkable domains, label-free alternatives such as majority voting or LLM-as-a-judge remove annotation cost but can introduce false positives that destabilize training. We introduce JURY-RL, a label-free RLVR framework that decouples answer proposal from reward disposal: votes from model rollouts propose a candidate answer, and a formal verifier determines whether that candidate can receive positive reward. Concretely, only rollouts matching the plurality-voted answer are rewarded when that answer is successfully verified in Lean. When verification is inconclusive, we invoke ResZero (Residual-Zero), a fallback reward that discards the unverified plurality proposal and redistributes a zero-mean, variance-preserving signal over the residual answers. This design maintains a stable optimization gradient without reinforcing unverifiable consensus. Across three backbone models trained on mathematical data, JURY-RL consistently outperforms other label-free baselines on mathematical reasoning benchmarks and transfers competitively to code generation and general benchmarks. It attains pass@1 performance comparable to supervised ground-truth training, with superior generalization demonstrated by higher pass@k and response diversity.

cs.AI

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

Chain-of-Thought (CoT) reasoning has become a powerful driver of trajectory prediction in VLA-based autonomous driving, yet its autoregressive nature imposes a latency cost that is prohibitive for real-time deployment. Latent CoT methods attempt to close this gap by compressing reasoning into continuous hidden states, but consistently fall short of their explicit counterparts. We suggest that this is due to purely linguistic latent representations compressing a symbolic abstraction of the world, rather than the causal dynamics that actually govern driving. Thus, we present OneVL (One-step latent reasoning and planning with Vision-Language explanations), a unified VLA and World Model framework that routes reasoning through compact latent tokens supervised by dual auxiliary decoders. Alongside a language decoder that reconstructs text CoT, we introduce a visual world model decoder that predicts future-frame tokens, forcing the latent space to internalize the causal dynamics of road geometry, agent motion, and environmental change. A three-stage training pipeline progressively aligns these latents with trajectory, language, and visual objectives, ensuring stable joint optimization. In inference, the auxiliary decoders are discarded, and all latent tokens are prefilled in a single parallel pass, matching the speed of answer-only prediction. Across four benchmarks, OneVL becomes the first latent CoT method to surpass explicit CoT, delivering superior accuracy at answer-only latency. These results show that with world model supervision, latent CoT produces more generalizable representations than verbose token-by-token reasoning. Code has been open-sourced to the community. Project Page: https://xiaomi-embodied-intelligence.github.io/OneVL

cs.CV

Enhanced Self-Supervised Multi-Image Super-Resolution for Camera Array Images

Conventional multi-image super-resolution (MISR) methods, such as burst and video SR, rely on sequential frames from a single camera. Consequently, they suffer from complex image degradation and severe occlusion, increasing the difficulty of accurate image restoration. In contrast, multi-aperture camera-array imaging captures spatially distributed views with sampling offsets forming a stable disk-like distribution, which enhances the non-redundancy of observed data. Existing MISR algorithms fail to fully exploit these unique properties. Supervised MISR methods tend to overfit the degradation patterns in training data, and current self-supervised learning (SSL) techniques struggle to recover fine-grained details. To address these issues, this paper thoroughly investigates the strengths, limitations and applicability boundaries of multi-image-to-single-image (Multi-to-Single) and multi-image-to-multi-image (Multi-to-Multi) SSL methods. We propose the Multi-to-Single-Guided Multi-to-Multi SSL framework that combines the advantages of Multi-to-Single and Multi-to-Multi to generate visually appealing and high-fidelity images rich in texture details. The Multi-to-Single-Guided Multi-to-Multi SSL framework provides a new paradigm for integrating deep neural network with classical physics-based variational methods. To enhance the ability of MISR network to recover high-frequency details from aliased artifacts, this paper proposes a novel camera-array SR network called dual Transformer suitable for SSL. Experiments on synthetic and real-world datasets demonstrate the superiority of the proposed method.

physics.optics

Perturb-and-Restore: Simulation-driven Structural Augmentation Framework for Imbalance Chromosomal Anomaly Detection

Detecting structural chromosomal abnormalities is crucial for accurate diagnosis and management of genetic disorders. However, collecting sufficient structural abnormality data is extremely challenging and costly in clinical practice, and not all abnormal types can be readily collected. As a result, deep learning approaches face significant performance degradation due to the severe imbalance and scarcity of abnormal chromosome data. To address this challenge, we propose a Perturb-and-Restore (P&R), a simulation-driven structural augmentation framework that effectively alleviates data imbalance in chromosome anomaly detection. The P&R framework comprises two key components: (1) Structure Perturbation and Restoration Simulation, which generates synthetic abnormal chromosomes by perturbing chromosomal banding patterns of normal chromosomes followed by a restoration diffusion network that reconstructs continuous chromosome content and edges, thus eliminating reliance on rare abnormal samples; and (2) Energy-guided Adaptive Sampling, an energy score-based online selection strategy that dynamically prioritizes high-quality synthetic samples by referencing the energy distribution of real samples. To evaluate our method, we construct a comprehensive structural anomaly dataset consisting of over 260,000 chromosome images, including 4,242 abnormal samples spanning 24 categories. Experimental results demonstrate that the P&R framework achieves state-of-the-art (SOTA) performance, surpassing existing methods with an average improvement of 8.92% in sensitivity, 8.89% in precision, and 13.79% in F1-score across all categories.

cs.CV

AGCD: Agent-Guided Cross-Modal Decoding for Weather Forecasting

Accurate weather forecasting is more than grid-wise regression: it must preserve coherent synoptic structures and physical consistency of meteorological fields, especially under autoregressive rollouts where small one-step errors can amplify into structural bias. Existing physics-priors approaches typically impose global, once-for-all constraints via architectures, regularization, or NWP coupling, offering limited state-adaptive and sample-specific controllability at deployment. To bridge this gap, we propose Agent-Guided Cross-modal Decoding (AGCD), a plug-and-play decoding-time prior-injection paradigm that derives state-conditioned physics-priors from the current multivariate atmosphere and injects them into forecasters in a controllable and reusable way. Specifically, We design a multi-agent meteorological narration pipeline to generate state-conditioned physics-priors, utilizing MLLMs to extract various meteorological elements effectively. To effectively apply the priors, AGCD further introduce cross-modal region interaction decoding that performs region-aware multi-scale tokenization and efficient physics-priors injection to refine visual features without changing the backbone interface. Experiments on WeatherBench demonstrate consistent gains for 6-hour forecasting across two resolutions (5.625 degree and 1.40625 degree) and diverse backbones (generic and weather-specialized), including strictly causal 48-hour autoregressive rollouts that reduce early-stage error accumulation and improve long-horizon stability.

cs.AI