Search arXivSearch

arXiv subjects

Yang Yu

Publications and source records attributed to Yang Yu.

At least 19 recordsLinked to original sources

Self-powered InAs nanowire detector arrays for extended-SWIR spectrometry at room temperature

Spectral sensing in the extended shortwave infrared (e-SWIR) is important for molecular analysis, infrared imaging, and machine vision, motivating the development of compact spectrometers for broader applications. However, conventional commercial off-the-shelf spectrometers in this wavelength region are expensive and bulky due to their reliance on external dispersive optics/filters and/or cryogenic accessories. Other emerging computational spectrometers are based on Si and InGaAs photodetectors that remain focused on the visible and near-infrared, with few detector platforms operating in the e-SWIR regime that simultaneously provide broadband sensitivity, low-noise room-temperature operation, and diverse spectral signatures for accurate identification and reconstruction. Here, we report a room-temperature e-SWIR computational spectrometer based on InAs/InP core-shell nanowire photodetector arrays with geometry-encoded spectral responses. The detectors exhibit self-powered broadband photoresponse across the 1--3 $\mu$m range, with responsivity up to 0.215 A W$^{-1}$, detectivity up to $1.6 \times 10^{9}$ cm Hz$^{1/2}$ W$^{-1}$, and microsecond response times. The excellent detector performance is leveraged to demonstrate filter-free spectral reconstruction using a compact multipixel photodetector array device. This enables high-accuracy molecular absorption spectrum reconstruction and hyperspectral imaging. Our results indicate that InAs nanowire arrays are a promising platform for compact computational spectrometry and imaging in the e-SWIR at room temperature.

physics.optics

Discrete Approximation to Time-changed Brownian Motions

We develop a general discrete approximation scheme for time-changed Brownian motions on $\mathbb{R}^d$. Our approximation scheme works for any smooth measure with full quasi-support on $\mathbb{R}^d$ with suitable initial distributions. Under some mild conditions on the smooth measure, the discrete approximation scheme works for every starting point. Our results in particular give a discrete approximation scheme for Liouville Brownian motions.

math.PR

R2S-Eval: Robot Evaluation with Real-to-Sim Calibration via Vision-Language Models

Evaluating robot manipulation policies is becoming increasingly important as generalist models, particularly vision-language-action (VLA) models, are deployed on physical robots. However, conventional real-world evaluation remains labor-intensive, unstable, and insufficiently informative. It requires repeated hardware trials, manual scene resets, and continuous operator monitoring, may produce different policy rankings across repeated evaluations, and primarily relies on success-rate metrics that provide limited information about execution quality. In contrast, humans assess robot performance by observing and comparing complete behaviors rather than relying solely on binary success outcomes. To this end, we propose R2S-Eval, an evaluation pipeline that combines real-to-sim calibration with vision-language model (VLM) preference evaluation. The real-to-sim component efficiently generates rollout videos in a simulator calibrated to the real-world evaluation setting, thereby reducing the need for repeated hardware trials. The VLM evaluator assesses the execution quality of rollout videos and produces pairwise preferences, which are subsequently aggregated into policy rankings. We further introduce a protocol to assess whether the proposed evaluation pipeline yields validated policy conclusions while mitigating the key challenges of conventional real-world evaluation. Experiments in both simulation and real-world settings demonstrate that R2S-Eval produces reliable and stable policy conclusions, achieves agreement with human preferences, substantially reduces repeated hardware-operation effort, and reveals behavior-quality differences that are not captured by binary success labels. In general, R2S-Eval advances robot evaluation from manual success counting toward automated, statistically stable, and quality-aware evaluation of robot behavior. Project page: https://r2s-eval.github.io.

cs.RO

Inversion Framework of Internal Mass Distribution Parameters of Asteroid Apophis from Dynamical Observations

The detection of the internal mass distribution of asteroids is of great significance for understanding their origin, evolution, and mission planning for exploration. Previous approaches rely on indirect density estimates or close-range spacecraft gravity inversion, which have limited applicability. This paper presents a proof-of-concept framework to infer the internal mass properties of asteroid (99942) Apophis during its close Earth flyby in 2029 using dynamical observations collected during the encounter. We establish a dynamical mapping from the evolution of orbital and rotational states to internal structural parameters, formulate it as an inverse problem, and solve it using Particle Swarm Optimization. The algorithm is first validated on a regular ellipsoidal model and then applied to three mass distribution models based on the actual shape of Apophis. Under ideal observation conditions, the relative error of the inverted moment of inertia ratios can be below 0.001%, and the absolute error of the center-of-mass position reaches the order of 10-5 meters. The algorithm successfully distinguishes among different internal structures. When realistic measurement noise is introduced, the inversion accuracy degrades. A sensitivity analysis reveals that the accuracy of the inertia tensor inversion is primarily limited by angular velocity measurement noise, whereas center-of-mass determination is highly sensitive to the precision of position and velocity data. This study provides a proof-of-concept for a technically feasible and cost-effective approach to infer asteroid internal structure during close encounters, and also highlights the critical data accuracy requirements for practical application, offering guidance for future observation campaigns.

astro-ph.EP

Atomic investigations on the mechanical properties of CoCrNi medium-entropy alloy nanowires

This work investigated the mechanical properties of CoCrNi medium-entropy alloy (MEA) nanowires under uniaxial tensile loading along three crystallographic orientations [100], [110] and [111]. The face-centered cubic (FCC) single-crystal CoCrNi nanowires exhibit pronounced anisotropy Molecular dynamics simulations were employed to provide atomic-level insights into the plastic deformation mechanism. The surface effect causes dislocations to nucleate at the free surface and slip inward. The temperature effect was also considered. Under identical geometric configurations, crystal orientations, and temperatures, interatomic interactions govern key mechanical properties such as Young's modulus and yield strength. Simulation results reveal a strong linear correlation between the average atomic force and mechanical performance across eight representative FCC metallic nanowires (Al, Au, Ag, Cu, Ni, Al19Mg alloy, FeCoCrCuNi high-entropy alloy (HEA), and CoCrNi MEA). This study delivers novel insights into the mechanical behavior of metallic nanowires and offers guidance for the rational design of alloy nanowires.

cond-mat.mtrl-sci

Searching for Solar-Basin Axionlike-Particle Decay with XMM-Newton Blank-Sky Observations

Axion-like particles (ALPs) bound in the solar gravitational field form the so-called ALP solar-basin. Since the two-photon decay of non-relativistic particles is approximately isotropic, this population can be searched for using observations in the anti-solar direction. In this work, we propose a search strategy for narrow decay-line signals from the ALP solar basin using \textit{XMM-Newton} blank-sky observations (XMM-BSOs) stacked spectra data taken in directions opposite to the Sun. By jointly fitting the signal and background model, we obtain limits on $g_{a\gamma\gamma}^{95}$ in the mass range $m_a=1.4\text{--}16~{\rm keV}$, with typical sensitivities of $g_{a\gamma\gamma}\sim10^{-10}\text{--}10^{-11}~{\rm GeV}^{-1}$. We have implemented the first anti-solar search for the solar basin, demonstrating that this strategy can exploit the stacked exposure of a large number of X-ray observations and provide a scalable analysis framework for future searches.

hep-ph

A particle method for the Boltzmann equation via amortized sampling from Green's function of the lifted linear operator

The collision operator for the Boltzmann equation is a nonlinear nonlocal operator. When lifted in the extended 2-particle space, it is viewed as the projection of a collisional linear operator. In this paper, we propose a particle method that samples the post-collision relative velocity directly from the Green's function of this operator (the transition probability of the generated time-continuous Markov chain). The normalizing flow amortized sampling is then proposed to reduce the sampling complexity. The resulted method takes $O(N)$ each time where $N$ is the particle number, and conserves momentum and energy exactly. This method does not require the boundedness of the kernel and, more importantly, it allows learning the Green's function directly from the scattering data without selecting the kernel in a specified family.

math.NA

ADE: Agentic Data Evolution Framework for Human-Centered Objectives

Aligning large language models to human-centered objectives is difficult when targets are non-executable and context-dependent, limiting reliable verification and scalable supervision. Although synthetic data expands coverage, weak verification shifts the bottleneck from generation to selection. Noisy signals destabilize iterative refinement and can cause silent regressions. We propose Agentic Data Evolution (ADE), a data-centric framework that organizes synthetic supervision as evolving data snapshots. ADE improves data snapshots through a closed-loop Observation-Variation-Selection (OVS) procedure, where a steady-state admission mechanism acts as a quality ratchet that conservatively gates updates for sustained cross-round improvement. We validate these improvements through complementary intrinsic trend tracking and extrinsic post-training evaluation. On DEV300, ADE raises the intrinsic win rate from 50% to 75.81% and the extrinsic win rate from 55.20% to 68.86%, consistent performance gains across diverse benchmarks. Blind expert evaluation further confirms this, with a 66.11% preference for evolved answers. These gains extend across post-training methods, model scales, and tasks beyond the target weakly verifiable educational objectives. Resources are available at https://github.com/ZeroLoss-Lab/Agentic-Data-Evolution.

cs.CL

On the Capability Separation Between World-Model Policy Learning and Imitated World-Action Models

World-action models predict a future outcome and then infer an associated action. Although this factorization can improve representation learning and data efficiency, it is unclear whether it provides stronger control capability than direct behavior cloning when both are trained from the same observational demonstrations. We compare a direct behavior-cloning policy, an imitation-trained world-action policy, and a policy optimized with an action-conditioned world model. At the controller-class level, every world-action policy can be flattened into a direct stochastic policy with the same closed-loop trajectory distribution. At the population level, under realizability, exact optimization, common deployment information, and distribution-preserving deployment, direct behavior cloning and world-action imitation both recover the observational behavior policy. Thus, future prediction changes the learning factorization but not the unrestricted external policy class or ideal imitation target. Action-conditioned world-model learning differs by predicting outcomes under specified actions and comparing them through a control objective. We characterize the irreducible action-specific prediction error of future models that do not condition on the candidate action, identify conditions under which a world-action joint can recover an interventional forward model, and show that observational demonstrations do not identify action effects in general. Finally, we construct an environment family in which every observational learner has positive worst-case regret, whereas one informative intervention permits zero regret. The key distinction is therefore between predicting futures associated with observed behavior and predicting consequences of specified actions for policy optimization.

cs.LG

UR$^{2}$-MLLM: Uncertainty-aware Revisit Reasoning in Multimodal Large Language Models for Radiology Report Generation

Radiologists generate diagnostic reports through iterative and selective revisiting of suspicious regions to refine their interpretations. Recent multimodal large language models (MLLMs) for radiology report generation (RRG) have shifted from text-only reasoning toward a ``Thinking-with-Images'' paradigm, incorporating visual evidence into the reasoning process. However, existing methods provide static visual evidence without a dynamic revisit mechanism during reasoning, neglecting how radiologists re-examine uncertain observations. To this end, we propose an Uncertainty-aware Revisit Reasoning MLLM (UR$^{2}$-MLLM) framework that dynamically revisits uncertain regions during reasoning for RRG. UR$^{2}$-MLLM is first equipped with uncertainty perception by training on an uncertainty-aware dataset. We then construct a multimodal reasoning trajectory dataset together with a detect-and-copy mechanism, which guides when and where to revisit. Finally, a visual grounding reward refines this behavior through reinforcement learning, aligning the revisited regions with corresponding anatomical structures. Experiments on MIMIC-CXR and IU-Xray show that UR$^{2}$-MLLM achieves state-of-the-art performance, highlighting the value of uncertainty-aware visual revisit reasoning for reliable and clinically aligned report generation.

cs.CV

From Association to Causation: Improving Retrieval Precision of Retrieval-Augmented Generation via Causal Relations and an Attention Mechanism

Retrieval-Augmented Generation (RAG) grounds LLM generation on retrieved documents, but the standard terminal retrieval stage--dense-vector similarity, optionally followed by reranking--often returns documents that share keywords with the query without containing the needed information, a failure mode that grows with the knowledge base. We trace it to a conceptual gap: similarity captures only associational relations, whereas the documents that matter are linked to the query causally. We model the terminal retrieval stage with a causal graph grounded in Reichenbach's common cause principle: the keywords shared by the query and a retrieved document form a latent common cause A, and the document's residual keywords form a latent set B linking the document to the ideal output. Since a retrieved document is a collider (A -> d <- B), retrieval itself opens an associational path between the query and B, which licenses a training-free, attention-style re-scoring rule: the cosine similarity between the query embedding and the weighted centroid embedding of B. Unlike causality-enhanced RAG variants that model causal relations inside the knowledge content, our graph models the causal structure of the retrieval process itself. On a real 471-document enterprise knowledge base, the method promotes a relevant guideline from rank 6 to the top 3; on a controlled diagnostic corpus reproducing the keyword-stuffing regime, it improves the mean target rank from 2.88 to 1.25, while a trained cross-encoder reranker barely helps (2.63). Conversely, on three BEIR benchmarks the score underperforms the similarity baseline, delineating the applicability boundary: the method guards the keyword-stuffing regime of growing proprietary knowledge bases and complements neural rerankers; a corpus-level calibration gate selects the correct regime with >= 95% reliability. A fully local testbed demonstrates deployability.

cs.AI

OpenHarmony Bench: Evaluating LLMs and Coding Agents on OpenHarmony App Development

We present OPENHARMONY BENCH, an app-level coding benchmark for evaluating LLM-based coding agents on OpenHarmony ArkTS applications. Unlike function-level benchmarks, it evaluates complete app-level changes: each task requires an agent to modify a buildable ArkTS project so that a requested behavior works end to end, involving UI state, data persistence, build configuration, and platform APIs. The benchmark installs and drives the delivered application on a device to check whether the behavior is observable. It covers three input sources: natural-language feature requests (new-feature), structured scenario specifications (spec-driven), and bug descriptions (bug-fix). The benchmark contains 153 top-level tasks and 242 Feature points (F-points), where an F-point is one executable behavior check. The snapshot includes 32 new-feature tasks, 50 spec-driven tasks with 139 F-points, and 71 bug-fix tasks. The main leaderboard is scored over top-level tasks rather than independently weighted F-points. We describe the benchmark construction, statistics, and build-and-test evaluation pipeline, and evaluate DevEco Code with eight LLMs across three independent full-suite runs per configuration. Three findings emerge. First, newer generations complete more tasks than their predecessors within evaluated model-family pairs. Second, buildability is close to saturated while behavioral correctness is not: mean Final Build Success Rate is 94.77% to 100.00%, whereas mean Task Completion is 48.36% to 58.39%. Third, spec-driven tasks have the lowest Task Completion under all-checks task scoring, with no configuration exceeding 35%. The code, data, tasks, reference solutions, tests, evaluation scripts, and leaderboard are released through the official OPENHARMONY BENCH website at https://bench.matrix.openharmony.cn/.

cs.SE

EA-LiteUNet: An Edge-Adaptive and Resource-Efficient U-Net for Boundary-Sensitive Dermoscopic Image Segmentation

Accurate boundary delineation remains a persistent challenge in dermoscopic image segmentation because of blurred lesion margins, heterogeneous textures, and complex background artifacts. From a signal-processing perspective, lesion boundaries represent high-frequency components that are highly susceptible to aliasing, noise amplification, and information loss. Consequently, repeated downsampling and feature transformations in conventional convolutional architectures often lead to severely degraded boundary representations. To address these limitations, we propose EA-LiteUNet, an edge-adaptive and computationally efficient U-Net variant specifically designed for boundary-sensitive medical image segmentation. The architecture integrates three core mechanisms: (1) boundary-aware representation learning to suppress aliasing and preserve high-frequency structural details; (2) attention-guided feature modulation to selectively enhance boundary-relevant responses across multi-scale features; and (3) a resource-adaptive inference strategy to dynamically balance segmentation accuracy and computational efficiency. Extensive evaluations across three public dermoscopic datasets demonstrate that EA-LiteUNet consistently achieves superior boundary precision. Specifically, on the ISIC 2018 dataset, the method significantly reduces the 95% Hausdorff Distance (HD95) to 12.89 pixels while maintaining a robust Dice score of 92.08%. Notably, this strong performance is achieved with an ultralightweight configuration of merely 0.29M parameters and 1.17 GFLOPs. Ablation studies further validate the complementary effects of these components, confirming their contribution to enhanced boundary fidelity and stable optimization.

cs.CV

A scalable edge-pass Purcell filter for high-fidelity readout of superconducting qubits

High-fidelity readout with strong Purcell protection of qubit coherence is essential for scalable superconducting quantum processors, yet the finite passband and sizable footprint of conventional band-pass Purcell filters make them hard to scale. Here we introduce a scalable edge-pass Purcell filter that separates the readout band from the protected qubit band by a single transmission edge, freeing the readout resonators from bandwidth constraint. Depending on whether the transmitting band lies above or below the cutoff, the compact network is realized as a high-pass filter (HPF) or a low-pass filter (LPF). The HPF reaches an average readout fidelity of 99.46(4)% (up to 99.56%) with a 150-ns pulse, and the LPF reaches 99.49(3)% (up to 99.57%) with a 130-ns pulse. The average single-qubit gate fidelities are 99.94% (HPF) and 99.93% (LPF). Relative to the filter-free Purcell limit, the filters substantially extend the qubit lifetime, and the Purcell protection deepens at higher filter order. In addition, an intrinsic dissipation mode of the filter offers a qubit-reset channel. This leads to a compact architecture that unifies fast, high-fidelity readout, Purcell protection, and effective reset within a single filter for large-scale fault-tolerant quantum computation.

quant-ph

Properties of holographic superconductors from Machine Learning

We investigate holographic superconductors using modern optimisation techniques inspired by machine learning. The critical temperature is obtained by minimising the variational functional for the eigenvalue $\lambda^2$ with two complementary trial functions: a simple cosine ansatz $F(z)=\cos(a z)$ and a flexible exponential polynomial $F(z)=\exp(\sum_{n=2}^{N} a_n z^n)$, both of which automatically satisfy the standard boundary conditions. For the cosine ansatz, we perform a one-parameter minimisation and obtain $\lambda^2(\Delta)$ and $T_c/\sqrt{\rho}$ over a wide range of $\Delta$, including the exact values at $\Delta=1$ and $\Delta=2$ to high accuracy. The exponential polynomial ansatz, with up to 19 coefficients, is optimised using a multi-start L-BFGS-B algorithm with warm-starting, yielding even better agreement with known exact results. Our numerical data for $\lambda^2(\Delta)$ and $T_c/\sqrt{\rho}$ match the analytical predictions from the literature, confirming the robustness of the variational approach. This work; therefore, demonstrates that a combination of analytic trial functions and modern numerical optimisation provides a powerful, flexible, and efficient tool for exploring holographic superconductors, and can be readily extended to include backreaction or other sectors in this field.

hep-th

Direct experimental measurement of femtonewton-scale momentum transfer force from electron beams

Electron beams (e-beams) are ubiquitous in imaging, patterning, and propulsion. This prevalence is rooted in the profound mastery of their wave-particle duality and energy-transfer pathways. Yet, a fundamental dimension remains largely unexplored: while the mechanical effect (i.e., the momentum transfer to a target) is theoretically known, quantification of its femtonewton-range force has remained elusive. This discrepancy represents a missing piece of the puzzle toward a comprehensive understanding of e-beam-matter interactions, and ultimately limits the multi-dimensional exploitation of e-beams. A force sensor combining femtonewton sensitivity, immunity to electromagnetic noise, compatibility with vacuum, and absolute calibration is critical to bridge the gap between theory and experiment. Here the FINEST (Femtonewton Interferometric Nanomechanical Electron-beam Sensing Technology) sensor is proposed and successfully tested to measure the force of an e-beam. FINEST is an optical-pressure-calibrated 3D spring-type optical sensor that operates reliably under e-beam conditions. Femtonewton-scale forces from 2-30 keV e-beams are directly measured, ranging from 505 fN to 13 pN. Both linear scaling with beam current and a non-monotonic energy dependence (peaking near 10 keV) are observed. Based on this calibrated force, the mechanical contribution to e-beam ice etching was quantitatively confirmed; its effect is orders of magnitude lower than the total etch depth and lacks noticeable energy dependence. By achieving the first direct experimental measurement of e-beam momentum transfer, this work adds a long-missing dimension to the physical landscape of e-beam processes. These findings provide a quantitative basis for furthering the multi-dimensional exploitation of e-beams, potentially transforming our approach to precision nanofabrication, sensing, and fundamental electron physics research.

physics.optics

ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is supposed to be accurate, safe under sensitive prompts, instructionally useful, and aligned with pedagogical goals at the same time. Existing benchmarks evaluate these requirements largely in isolation, so none assesses education-facing suitability as an integrated profile. We introduce ELBench, the first benchmark to evaluate all four requirements (General Capability, Safety and Trustworthiness, Basic Education, and High-Level Cultivation) on the same models under a common protocol, combining curated public sources with newly synthesized safety and cultivation data. We evaluate nine models, seven frontier general-purpose systems and two education-specialized variants, and report three findings. First, module-level profiles are more informative than a single aggregate: the top six models are statistically indistinguishable on overall score, yet their module leaders differ substantially, and safety is anti-correlated with practical teaching (r = -0.83). Second, the Chinese-developed models lead the safety module, the most discriminative in the suite; this advantage is largest on region-specific normative content and narrows, but does not vanish, on universal-harm content. Third, the two education-specialized models lead neither education module, and on High-Level Cultivation all models share a systematic blind spot: on the structured judgment task they converge on the same non-reference option, favoring pedagogical style over fit to the stated goal, so the module scores uniformly low and does not separate models. This raises, but does not resolve, whether domain post-training keeps pace with frontier systems on education tasks.

cs.CL

Potential Matching Optimal Transport: Continuous Normalizing Flows for Exact $p$-Wasserstein Dynamics

We introduce Potential Matching Optimal Transport (PMOT), a potential-flow framework for general $p$-cost optimal transport with $c_p(x,y)=\|x-y\|^p$. PMOT parameterizes the CNF velocity field with a scalar potential in the generalized Benamou--Brenier form for the chosen exponent $p$. It trains the potential gradient with a self-induced matching loss along straight bridges determined by the model's own endpoints, while allowing flexible terminal distribution matching. Our main result establishes zero-loss exactness: under the stated regularity, exact terminal matching, and uniqueness assumptions, any zero-loss solution satisfies the generalized Benamou--Brenier optimality system and recovers the corresponding $p$-optimal transport map and dynamics. On synthetic benchmarks, PMOT learns $p$-specific maps that agree with the corresponding $p$-matched OT references. It also remains competitive as a likelihood-based density model on high-dimensional tabular data, and an MMD-based color transformation experiment demonstrates flexible sample-based terminal matching.

cs.LG