Search arXiv⌕ Search

arXiv subjects

Jie Yang

Publications and source records attributed to Jie Yang.

At least 19 recordsLinked to original sources

TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent

Time series agents answer analytical questions by calling external tools, and which tools they carry is decided by people before the agent runs. However, we identify two failures in this setup. Human-Agent Tool Misalignment: a library of 21 expert-curated tools helps on some tasks and hurts on others, dropping anomaly accuracy under every backbone we test. Silent Harm: one round of generic self-revision changes 147 answers and breaks 56 of them, while the final score moves by less than a point. Both follow from the same gap: whether a tool helps is decided question by question at runtime, while tools are supplied in advance and judged by a single average. To address this, we propose TimeEvo, which clusters an agent's diagnosed failures into capability gaps, plans a measurement for each, synthesizes evidence-only tools that fill them, and admits the candidate library only through a paired admission gate. Experiments on ten time series QA tasks and three backbones show that TimeEvo, starting from an empty library, improves accuracy on every task and every backbone, and that a library grown on a cheap model still gains when it is installed into stronger ones. Code is available at https://github.com/Muyiiiii/TimeEvo.

cs.AI↗

LastOPD: Taming Collapse in Latent On-Policy Distillation

On-policy distillation (OPD) corrects a student on the responses it writes, but its signal is the teacher's next-token distribution: it tells the student what the teacher says but misses how it thinks. Latent supervision promises the missing part by aligning the student's latent states to the teacher's. Recent methods such as OPRD bring this signal into on-policy distillation. However, we observe two failures of this recipe when distilling Qwen3-4B and Qwen3-8B into Qwen3-1.7B-Base. Early gain, late collapse: latent supervision alone lifts MATH-500 accuracy from 25 to 46 in 10 steps, but subsequent training degrades performance down to 11 with no recovery. Better alignment, worse behavior: although the alignment metric steadily improves throughout this collapse, the most aligned model turns out to be the worst performing. Further analysis suggests a mismatch in how the latent signal is applied: layers paired by depth play different roles in the two models, so continued alignment may pull the student toward teacher states it cannot understand. To address this, we propose LastOPD, which applies the latent signal only at the last-layer state, the common interface both LM heads read, and only during a 10-step crossfade into token-level OPD. This keeps the useful part of the latent signal and hands the student to token-level supervision before the collapse sets in. Extensive experiments show that LastOPD improves MATH-500 over token-only OPD by 5.55 and 4.02 points with the 4B and 8B teachers, leads on most held-out datasets, and reaches the final score of token-only OPD in about half the steps. Code is available at https://github.com/Muyiiiii/LastOPD.

cs.LG↗

MemAudit: Auditing Long-Term Agent Memory via Hidden User-State Recovery

Long-term memory promises LLM agents that grow more capable across sessions, maintaining an accurate, evolving understanding of the user that interaction forms. In practice, however, this memory is evaluated mostly through downstream behavior, such as later answers, personalization quality, or task success, which tests that understanding only indirectly and leaves the memory artifact itself largely unaudited. We argue that long-term memory should instead be evaluated as an auditable post-interaction artifact: after ordinary assistance, what structured user state can be reconstructed from the memory the agent leaves behind? We instantiate this view in MEMPROBE, a benchmark in which a memory-equipped agent assists simulated users, each carrying a hidden, taxonomy-anchored user-state bank, across a trajectory of leak-controlled tasks, after which that bank is reconstructed from the agent's resulting memory under both full-store and top-k access. Built on synthetic ground truth for efficient, scalable measurement, MEMPROBE spans 50 simulated users with 31 hidden dimensions each (1,550 recovery targets) and tests 5 representative memory systems. Testing state-of-the-art memory agents, we find that successful assistance and recoverable memory behave as distinct capabilities. Task completion nearly saturates, even for a memoryless baseline, while category-balanced recovery stays moderate (about 0.6) and drops further under top-k retrieval. MEMPROBE is the first benchmark to study memory recovery directly, reconstructing the user state a system retains and scoring it against ground truth. We see recovery as a concrete objective for future memory agents to optimize, and MEMPROBE as a step toward an environment where agents are trained to remember their users, growing more faithful the longer they know them.

cs.CL↗

DiaVLo: Diagnosing Behaviours of Vision-Language Models

Vision-language models (VLMs) rely on storing and transferring appropriate information across their sub-components. Verifying that the VLMs exhibit desired behaviours, while avoiding harmful ones, is central to their reliable deployment. Yet, methods that identify VLM behaviours remain scarce. We present DiaVLo, a diagnostic framework that leverages human curation and VLMs' generation capabilities to construct specifications of desired and observed VLM behaviours, surfacing potential misalignments. Beyond this, DiaVLo also provides causal estimates to identify the most influential concepts steering VLM behaviours. We evaluate DiaVLo on several open-source VLMs under both classification and generation conditions. Our experiments show that DiaVLo produces behaviour labels that correlate with model performance and provide context for measured performance. DiaVLo surfaced behaviours that are clearly aligned and misaligned, alongside patterns in how VLMs perceive, organise, and prioritise concepts.

cs.CL↗

Environment-Aware Diffusion Model for Massive MIMO-OFDM Channel Estimation

This paper proposes an environment-aware diffusion based channel estimation in massive multiple-input multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) systems. The high dimensionality of massive MIMO channels combined with limited pilot resources makes accurate estimation challenging. To address this issue, we exploit the spatial variability of wireless channels by training a diffusion model to learn the location-conditioned distribution of channel state information, which provides an environment-aware prior for channel estimation. Based on this learned prior, a posterior inference algorithm is developed to incorporate pilot observations into the reverse diffusion process, enabling Bayesian channel estimation by combining the received-signal likelihood with the learned channel prior. By jointly leveraging location information and measurement data, the proposed approach improves estimation accuracy under limited pilot resources. Simulation results based on ray-tracing channel datasets demonstrate that the proposed method consistently outperforms conventional estimators and existing learning-based approaches across various signal-to-noise ratios and pilot configurations.

eess.SP↗

Ambiguity Function Analysis of OFDM Signals With Pilots and Data Payloads

Practical orthogonal frequency division multiplexing (OFDM) communication frames contain both deterministic pilots and random data payloads, motivating the joint ambiguity function (AF) analysis of the two components when the entire frame is reused for integrated sensing and communication (ISAC). This paper characterizes two discrete AF formulations for different Doppler regimes, namely the discrete periodic AF (DP-AF) and fast-slow-time AF (FST-AF), and derives closed-form expressions for their expected squared values. For the FST-AF, the expected sidelobe level (ESL) is uniform over the delay-Doppler plane and depends only on the pilot count, constellation kurtosis and total number of time-frequency resources, but not on the pilot symbols or pattern. For the DP-AF, we establish attainable lower and upper ESL bounds and show that no pilot design can minimize all sidelobes simultaneously. We further prove that attaining the lower bound at non-zero Doppler requires a periodic pilot pattern, while equally spaced chirp pilots, including Zadoff-Chu (ZC) sequences, maximize the numbers of sidelobes attaining the lower and upper bounds simultaneously. Two representative ZC pilot patterns widely encountered in communication frames are then examined: contiguous placement produces delay-Doppler ridges described by squared Dirichlet kernels, whereas equally spaced placement generates periodic peak-and-notch structures. Both regular patterns exhibit pronounced high sidelobes, suggesting that communication-oriented pilot patterns should be re-designed for delay-Doppler estimation in the context of ISAC. Numerical results validate the analysis and show that irregular pilot placement can suppress high sidelobes and improve target estimation performance.

eess.SP↗

HLC-GS: Risk-Map-Guided Height-Layer Consistency Gaussian Splatting for DSM Reconstruction from Optical Satellite Imagery

A Digital Surface Model (DSM) is a fundamental geospatial data product for representing the elevation of the Earth's surface. Recently, 3D Gaussian Splatting (3DGS) has shown considerable potential for DSM reconstruction from multi-view optical satellite imagery due to its explicit scene representation and efficient optimization. However, in 3DGS-based DSM generation, alpha-weighted aggregation of Gaussian altitudes may blend splats from different height layers at the same rendered pixel or DSM sampling location, producing non-physical intermediate elevations and height-layer mixing errors. To address this problem, we propose HLC-GS, a risk-map-guided height-layer consistency Gaussian Splatting method for DSM reconstruction from optical satellite imagery. HLC-GS consists of a risk map module, a dominant-layer reliability correction module, and a secondary-layer suppression module. The risk map localizes high-risk pixels with abnormal height dispersion and unreliable dominant-layer responses, while the latter two modules regularize unreliable dominant-layer responses and suppress weakly supported far secondary-layer responses. Extensive experiments are conducted on the DFC2019 and IARPA2016 datasets. Compared with six state-of-the-art DSM reconstruction methods, HLC-GS achieves better overall accuracy. Compared with the latest and precision-enhanced EOGS, HLC-GS reduces the average MAE from 1.46 m to 1.18 m and the average RMSE from 2.78 m to 2.58 m over the evaluated scenes, while improving PAG$_{2.5}$ from 86.09\% to 88.61\%. Overall, these results demonstrate that explicitly modeling per-pixel height-layer consistency alleviates height-layer mixing and improves the geometric quality of 3DGS-based DSM reconstruction from optical satellite imagery.

cs.CV↗

Relations between the higher Hamiltonians of the trigonometric and the rational spin Calogero-Sutherland models

In this paper we study the relations between the Hamiltonian hierarchy generated by trigonometric Cherednik-Dunkl operators and the one generated by rational Dunkl operators. We develop two nested structures relating these two hierarchies. The first nested relation reconstructs the higher rational spin Calogero-Sutherland Hamiltonians exactly from the trigonometric ones. The second nested relation reconstructs the higher trigonometric spin Calogero-Sutherland Hamiltonians as leading terms from the rational ones. In addition, we provide an explanation of the eigenvalues of the trigonometric spin Calogero-Sutherland Hamiltonians in terms of $N$-colored Young diagrams and Maya diagrams.

hep-th↗

The 2PN Point-Mass N-Body Equations of Motion in Harmonic Gauge: A Computable Formulation

We develop a semi-analytic and semi-numerical formulation of the harmonic-gauge second post-Newtonian (2PN) equations of motion for a general point-mass N-body system within Hadamard regularization. The equations of motion are separated into a closed analytic contribution and a non-closed integral contribution. We analyze the singular structure of the latter and further regularize it into a numerically evaluable representation. We apply the formulation to the Sun-Jupiter-Saturn and Sun-Mercury-Venus systems, evaluating the instantaneous non-closed 2PN acceleration along Newtonian trajectories and its leading finite-time relative-distance response through the corresponding perturbation equations. In both benchmarks, the non-closed acceleration remains a small fraction of the complete 2PN acceleration, while the induced relative-distance perturbation remains oscillatory and can reach larger oscillation amplitudes at later times.

gr-qc↗

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation

We present Visko Orbis 1.0, a Live Model for real-time, interactive long video generation. Users can change the prompt at any moment during generation, and the update becomes visible in real time. Visko Orbis 1.0 supports long-form text-to-video, image-to-video, and video continuation, with multilingual prompts and prompt switching while generation is in progress. A bounded multi-scale memory preserves subjects, scenes, and style across chunks, sustaining hour-scale rollouts without evident quality or color drift. The generator is factorized causally in time, matching the causal structure of physical dynamics, and is aligned with a latent world-model reward for predictive consistency. Built on a distilled chunk-wise streaming generator and a streaming video upscaler, Visko Orbis 1.0 delivers 4K video generation at 24 FPS in real time, using an optimized GPU serving engine. In quantitative evaluations, Visko Orbis 1.0 achieves the best DOVER aesthetic and technical scores and the best VideoAlign visual and motion quality, and leads three physical-plausibility protocols (VideoPhy-2, Physics-IQ, and VBench-2.0 Physics); in long-form Arena comparisons, it obtains the highest overall-preference and temporal-stability ratings among all the state-of-the-art real-time interactive video generation systems.

cs.CV↗

Integrated yoctosecond-precision timing detector

Precise timing detection is essential for exploring ultrafast phenomena in fields ranging from free-electron lasers to ultra-high-power laser facilities. However, achieving simultaneous high resolution and large dynamic range remains a fundamental challenge, and state of the art systems are often constrained by their physical size, power requirements, and limited scalability. Here we introduce an integrated dual electro optic sub cycle timing detector (DEST) that overcomes these limitations. In a proof of principle measurement, the device resolves timing jitter as small as 11 yoctoseconds (ys, $10^{-24}$ s) at 1 MHz--equivalent to the transit time of light across two protons--while maintaining an unambiguous measurement range of 6.15 ps and a dynamic range exceeding 155 dB. The core detection unit is miniaturized to chip scale dimensions of 18 mm x 2 mm x 1 mm on a thin film lithium niobate platform, ensuring inherent stability and immunity to environmental disturbances. Moreover, the architecture naturally lends itself to massive parallelization through array integration, with the potential to push timing precision to the sub 10 rontosecond (rs, $10^{-27}$ s) level within a 1 $m^2$ footprint. This combination of extreme sensitivity, wide dynamic range, compact size, and scalability opens new avenues for detecting previously inaccessible weak signals, including those from gravitational waves, quantum vacuum fluctuations, and beyond.

physics.optics↗

ForLion: An R Package for Finding Optimal Experimental Designs with Mixed Factors

Optimal design is crucial for experimenters to maximize the information collected from experiments and estimate the model parameters most accurately. ForLion algorithms have been proposed to find D-optimal designs for experiments with mixed types of factors. In this paper, we introduce the ForLion package which implements the ForLion algorithm to construct locally D-optimal designs and the Expected Weighted (EW) ForLion algorithm to generate robust EW D-optimal designs, which maximize the determinant of the expected Fisher information matrix under parameter uncertainty. The package supports experiments under linear models (LM), generalized linear models (GLM), and multinomial logistic models (MLM) with continuous, discrete, or mixed-type factors. It provides both optimal approximate designs and an efficient function converting approximate designs into exact designs with integer-valued allocations of experimental units. Tutorials are included to show the package's usage across different scenarios.

stat.CO↗

Flow3D-OPD: Multi-Teacher On-Policy Distillation for 3D Geometry Generation with Flow-Matching Diffusion Transformer

Recent image-to-3D generation models built on flow-matching diffusion Transformers (DiT) can produce high-fidelity meshes, yet their post-training strategy remains largely unexplored. There exist several critical bottlenecks in reinforcement learning: the inherent difficulty of defining comprehensive rewards for 3D geometric quality, and the gradient interference that arises when jointly optimizing heterogeneous objectives. Inspired by the practicability of on-policy distillation (OPD) in large language models and image generation, we propose \textbf{Flow3D-OPD}, a two-stage post-training framework that introduces multi-teacher distillation into 3D geometry generation. In the first stage, we utilize the semi-policy to enhance the foundational capability of the pretrained model and then design an agentic verifier for 3D geometric quality evaluation. Based on the verifier, we could cultivate domain-specialized teacher models via direct preference optimization (DPO). In the second stage, we consolidate heterogeneous expertise into a unified student model through on-policy distillation with hard task-routing sampling and gradient accumulation, which could mitigate the gradient interference in joint optimization. Without relying on elaborate modifications, our straightforward yet effective design achieves consistent improvements across all geometric quality dimensions and surpasses all teacher models in the average metric. Extensive experiments demonstrate that our approach provides an effective paradigm for reinforcement learning in 3D generation.

cs.CV↗

Nonlinear path-following via the asymptotic numerical method on a quantum processor

Quantum computing offers a promising avenue for advancing computational methods in science and engineering. In this work, we introduce the quantum asymptotic numerical method (qANM), a framework for solving nonlinear path-following problems using quantum computing. Based on the principle of high-order perturbation techniques, the proposed method uses Taylor series expansions to transform complex nonlinear systems into sequences of linear equations, which are then solved using quantum linear solvers. The central objective of this study is to demonstrate nonlinear path-following with the required linear systems solved on real quantum hardware. To realize this objective, we develop the quantum-enhanced Jacobi method (q-Jacobi), an iterative quantum linear solver used in the hardware experiment. Numerical simulations on a quantum simulator validate the convergence of the method. A highlight of this work is a proof-of-principle experiment on a superconducting quantum processor. Despite the noise inherent in near-term quantum hardware, the experiment achieves 98% accuracy in tracking the nonlinear solution path. We believe this work provides a useful reference for applying quantum computing to nonlinear computational mechanics.

quant-ph↗

PAI-Actor: Cinematic Multi-Character Replacement in Dynamic Scenes

We present PAI-Actor, a cinematic multi-character animation framework for character replacement in dynamic movie scenes. Unlike conventional animation systems that mainly drive a single static image or a single subject, our goal is to replace and animate multiple characters within real video clips while preserving the original scene dynamics, camera motion, and background content. This setting is particularly challenging because the generated characters must remain consistent with the source performance in motion and interaction, while also matching the surrounding background in lighting, shadow, composition, and overall cinematic appearance. To address this, we formulate multi-character animation as a structure-guided human recovery problem and build a movie-driven training pipeline from high-quality film data. Furthermore, to support practical cinematic production, we introduce a bidirectional-to-autoregressive distillation framework: we first train a bidirectional diffusion transformer for high-quality short-clip generation at 1080P resolution, and then distill it into an autoregressive video-to-video model for efficient inference and longer video generation. Experiments show that PAI-Actor enables high-fidelity multi-character animation with strong scene consistency, cinematic visual quality, and efficient long-form generation.

cs.CV↗

EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents

Computer-use agents must solve long-horizon tasks through repeated interaction with partially observable, multimodal desktop environments. Although imitation learning and offline trajectory refinement provide strong priors, static traces cannot cover the causal feedback loop of real computer use: each action changes the screen state, future action space, and recovery options. EvoCUA-1.5 extends self-evolving computer-use agents from offline experience learning to online reinforcement learning, where policies interact with executable sandbox environments and improve from verifiable task outcomes. Online RL in this setting requires more than directly reusing single-turn language-RL recipes. Multi-turn interaction introduces context-managed observations, sparse terminal rewards, variable-length trajectories, and slow environment feedback. EvoCUA-1.5 addresses these challenges with Step-Level Policy Optimization (STEPO), which preserves trajectory-level advantage balance after decomposition into step-level samples; policy-aware filtering and pass-rate calibration over verifiable synthesized tasks; Dynamic Tri-Adaptive Curriculum (DTAC), which combines learnable tasks, difficult positive replay, and controlled infeasible-task exposure; and a fully asynchronous RL infrastructure with staleness control and mini-group batching. Experiments show that these components improve training stability and downstream performance. EvoCUA-1.5 achieves 63.2\% success on OSWorld-Verified, outperforming comparable 32B/35B-scale open-weight baselines and even approaching models with significantly larger parameter counts. Overall, EvoCUA-1.5 provides a practical framework for scaling online RL in multi-turn computer-use agents.

cs.AI↗

EAVer: Long-Form Factuality Verification as an End-to-End Agentic Policy

Long-form factuality verification is commonly implemented as a static decompose-search-verify pipeline, with separately prompted modules processing claims and invoking external search. Treating claims independently makes LLM and search calls scale with claim count and causes repeated searches for overlapping evidence about related claims. We introduce EAVer, an End-to-end Agentic Verifier that learns to control the complete response-level verification workflow as a unified policy. EAVer groups semantically related claims, routes each group to direct verification or targeted search based on confidence, and keeps evidence returned by search in compact in-context memos for cross-claim reuse. To train this policy, we develop a privileged-teacher synthesis pipeline that converts gold claim annotations into executable multi-turn tool-interaction trajectories with live search rather than post-hoc rationales. Structural, label-alignment, tool-use, search-budget, and leakage checks yield 1,447 quality-controlled trajectories. We further construct 794 bidirectional same-trajectory preference pairs that keep claim grouping, search, and evidence fixed, enabling decision-focused Direct Preference Optimization (DPO) over factuality-decision tokens. The results with Qwen3-8B show that EAVer outperforms the strongest search-based baseline on each benchmark by 2.88 Macro-F1 points on VeriFastScore and 4.73 points on the out-of-distribution FaStFact-Bench, while using about 80% fewer searches than the most search-efficient baseline. Moreover, EAVer consistently improves performance across models ranging from 4B to 32B parameters, demonstrating its strong generalizability.

cs.CL↗

Muonium Spectroscopy as a Quantum Sensor for Ultralight Axion Dark Matter

High-intensity muon beams could enable a Muonium-based Axion Search through resonant quantum transitions between Hyperfine states (MASH). Combining theoretical calculations with simulation results, we demonstrate that such a muonium-based experimental approach could complement and tighten constraints on the axion-muon coupling beyond existing limits from the muon $g\!-\!2$ measurement, over the axion mass range of 18--130 $μ$eV. These results establish a new spectroscopic channel in muonium that enables searches for axion and axion-like particle dark matter.

hep-ph↗