Search arXivSearch

arXiv subjects

Ting Li

Publications and source records attributed to Ting Li.

At least 19 recordsLinked to original sources

Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents

Self-evolving runtime harnesses can substantially improve the capabilities of large language model (LLM) agents and provide a promising paradigm for optimizing agent execution. Existing harness evolution methods typically rely on iterative search, repeatedly evaluating and revising candidate harnesses based on execution feedback from task instances. While this paradigm enables continuous harness optimization, it incurs substantial time overhead due to repeated agent executions and code modifications, and may overfit to observed tasks and specific failure patterns, resulting in degraded generalization to unseen tasks. We identify the lack of principled failure diagnosis as a key bottleneck in harness evolution: an observed failure can reflect either model-specific deficiencies or systematic harness deficiencies, and directly optimizing against individual failures can lead to unnecessary model-specific accommodation. We therefore propose Ecdysis, an efficient and effective framework that distinguishes model-specific accommodation from harness-level repair and biases adaptation toward systematic harness deficiencies by identifying recurring cross-task failure patterns. Ecdysis adopts a batch-level cross-instance failure aggregation paradigm to jointly analyze failure evidence from multiple task instances and further introduces Failure-Driven Collaborative Refinement to diagnose failure causes and iteratively refine harness modification specifications. By combining cross-instance failure analysis with multi-role diagnosis, Ecdysis enables more effective harness evolution with lower training time. Experiments show that Ecdysis achieves up to a 1.84x speedup in harness training compared with existing harness evolution methods, while improving the reasoning accuracy of the resulting harnesses by 18.56%.

cs.SE

Empirical-Bayes Elastic-Net Computation for Exponential Random Graph Models

Exponential random graph models (ERGMs) describe dependence among network ties, but inference becomes difficult when the likelihood is intractable and candidate network statistics are strongly correlated. We introduce BERGM Elastic Net, an adaptive empirical-Bayes approach that combines lasso shrinkage with ridge stabilization in a Bayesian ERGM. A latent-variable formulation supports approximate exchange sampling, while empirical-Bayes updates adapt the amount of regularization to the observed network. We connect the proposed prior to elastic-net penalized likelihood and clarify the interpretation of thresholded reporting and coefficient grouping. The method is developed for over-specified network models containing many related structural and covariate effects.

stat.ME

Statistical Study of Solar Prominence Plumes Based on NVST H$α$ Observations

Plumes are one of the most representative dynamic features observed in prominences and play a key role in mass and magnetic transport within them. However, their physical nature and triggering processes remain actively debated. Based on limb H$α$ observations from the New Vacuum Solar Telescope (NVST) during 2013--2025, we statistically investigated 34 plumes with clear and complete evolutions by developing an automated image-processing pipeline. It is revealed that plume lifetimes mainly range from 300 s to 700 s, with vertical displacements between 3--7 Mm. The mean widths and velocities are concentrated in the range of 0.5--1.5 Mm and 10--20 km s$^{-1}$, respectively. Besides wide distribution ranges, plume parameters exhibit irregular evolution fluctuations, indicating that the formation and evolution of various plumes may exhibit different physical patterns. Correlation analysis among the parameters further reveals that: (1) Positive correlations were found among lifetime, vertical displacement, and mean width, indicating an intrinsic coupling between the temporal and spatial scales of plumes. (2) Trajectory curvature is negatively correlated with lifetime, vertical displacement, and velocity. Accelerating and width-contracting plumes typically have lower curvature, suggesting that curvature may reflect environmental influences and the stability of plumes. (3) Plumes with higher initial velocities were more likely to be accompanied by precursor brightening, suggesting that these plumes may be triggered by magnetic reconnection. Furthermore, we infer that some plumes in non-bubble regions may be inherently driven by mini-filament eruptions. These results establish a statistical framework for prominence plumes and reveal diversity in their dynamical evolution and triggering mechanisms.

astro-ph.SR

Offline-Online Curriculum RL for Multimodal Reasoning

Multimodal large language models exhibit capabilities on reasoning tasks, yet often produce flawed intermediate steps while yielding correct final answers. This behavior undermines interpretability and reliability, suggesting reliance on spurious shortcuts rather than faithful reasoning. Although efforts have explored step-level supervision, distinguishing decisive steps from redundant ones remains challenging. We propose $O^2$-CritiCuRL, a novel curriculum reinforcement learning framework that introduces critical-step awareness through an iterative offline-online paradigm. In the offline stage, $O^2$-CritiCuRL conducts multi-rollout analysis over step-annotated trajectories to estimate step-level importance, allowing the framework to distill critical reasoning steps and filter out redundant ones. In the online stage, we employ a progressive step-level reinforcement learning strategy, where truncated chains guide the model to infer missing steps and refine its reasoning, thereby sharpening its focus on critical steps and overcoming the limitations of static supervision. Extensive experiments on multimodal reasoning benchmarks show that our method achieves state-of-the-art performance while delivering superior training and inference efficiency. Code is available at https://github.com/kk0013/CritiCuRL.

cs.AI

Efficient Online LLM Watermark Detection via Rao-Blackwellized E-Processes

As large language models (LLMs) are increasingly deployed, reliable and efficient mechanisms for distinguishing AI-generated text from human-written content have become essential. Statistical watermarking has emerged as a promising solution, yet most existing methods are typically fixed-horizon procedures, precluding valid early stopping in streaming generation. In this paper, we develop an efficient online watermark detection framework with anytime-valid inference based on Rao-Blackwellized e-processes, enabling recursive token-level evidence updates without storing the full history. In particular, we instantiate the framework for the Gumbel-max watermark and reduce the original token-level dependence testing problem to a pivot-induced sequential testing problem with an explicit null distribution. Theoretically, we prove anytime-valid Type I error control under arbitrary optional stopping and establish positive asymptotic log-growth under watermarking, implying consistency of the proposed stopping rules. Simulations and experiments on real LLM-generated text demonstrate efficient online detection with rigorous anytime-valid guarantees.

stat.ML

Coarse-to-Fine: A Hybrid Self-Supervised Method for Non-rigid 3D Shape Matching

Non-rigid 3D shape matching is a fundamental task in computer vision and graphics. In this paper, we propose a hybrid self-supervised method based on a coarse-to-fine strategy, which ensures consistency between the coarse mapping and the refined correspondence produced by our refinement module. The architecture features a dual-branch design, consisting of two symmetric functional map learning streams: one based on the Laplacian basis and the other utilizing the elastic basis. Extensive experiments show that our approach not only maintains computational efficiency, but also achieves state-of-the-art performance across a variety of challenging scenarios, including non-isometric deformations and topological noise. Finally, we rigorously demonstrate that contrastive energies promote feature discrimination. Furthermore, integrating these energies with existing methods yields consistent improvements, validating the overall efficacy of our approach. Our code is available at https://github.com/LuoFeifan77/Coarse-to-Fine-Hybrid-Self-Supervised-Matching.

cs.CV

Unprecedent fast winking of solar flares triggered by bursty magnetic reconnection

Flare ribbons form as a result of energy deposition associated with particles accelerated in low layers of the solar atmosphere. The fine-scale structures of flare ribbons, also called ribbon kernels, offer a potentially powerful diagnostic of the flare reconnection process, however to date the dynamic evolution of ribbon kernels has not been fully characterized in statistical studies. Here, we checked the state-of-the-art observations (cadence $\leq$ 2.5 seconds) of solar flares in the ultraviolet from space by Interface Region Imaging Spectrograph (IRIS) over the past 12 years. Our results showed the first statistical study of multiple spatially-resolved flare kernel quasi-periodic pulsation events for 31 flares, with the period of 6-24 seconds. The ribbon kernels have a spatial scale of 480$-$1200 km and some kernels exhibit unprecedent fast ``winking" process, i.e., quasi-periodic pulsation-like flashing of individual kernels. The shortest heating time reaches about 2$-$3 s, implying that the energy is deposited only in a small localized region within flare ribbons, persisting for only a few seconds. Meanwhile, some ribbon kernels were observed to slip along the ribbon at speeds of 20-1800 km s$^{-1}$. These observations strongly imply a joint picture for the dynamics and the bursty nature of ribbon kernels as being due to coupled effects of plasmoid formation and three-dimensional (3D) magnetic reconnection in the overlaying coronal current sheet. We suggest that the observed flare behaviors provide strong observational evidences of 3D bursty reconnection.

astro-ph.SR

Revisiting Approaches to Stellar White-Light Flare Energy Based on Spatiotemporally Resolved Solar Observations

Accurately estimating the bolometric energy of solar and stellar white-light flares (WLFs) is crucial for understanding their physical nature and impact on surrounding planets. However, the lack of spatial resolution in stellar observations forced pioneering stellar WLF studies to adopt simplified energy estimation methods, typically assuming either a constant flare temperature or a fixed radiating area. To assess the physical plausibility of these assumptions, we utilize high-spatiotemporal-resolution solar observations to analyze the true evolution of source region's radiating area and temperature of 70 solar WLFs. It is revealed that both area and temperature of most solar WLFs undergo significant temporal evolution, and the flare area strongly correlates with the flare's peak optical continuum flux. Therefore, we propose a new energy estimation method that permits both flare area and temperature to evolve. Compared with existing methods, our dynamic approach yields systematically lower flare energies, which then prompts us to revisit classical macroscopic scaling laws related to the flare energy. It is further revealed that different energy estimation approaches can systematically alter these scaling relations, calling for a re-examination of these established statistical results and their targeted testing or revision in future work.

astro-ph.SR

Global Average Treatment Effects for Individualized Randomization Experiments with Aggregate Data

Individualized randomized experiments are central to online platforms for optimizing personalized decisions in complex environments. In two-sided markets, however, standard treatment effect estimation is often invalid due to strong temporal and cross-unit interference, a challenge compounded when only aggregated data are available because of privacy or system constraints. To address these issues, we identify the Global Average Treatment Effect (GATE) using only group-level data from treatment and control groups. We first establish identification conditions based on aggregated observations, and then propose the Individualized Randomized Experiment Varying Coefficient Decision Process (IRE-VCDP) model, which accounts for interference through supply-demand dynamics. Building on this framework, we develop a complete procedure for estimation and statistical inference of the GATE, along with theoretical guarantees for the proposed test. Extensive simulations and real-world experiments using data from a leading ridesharing platform demonstrate the effectiveness of our approach.

stat.ME

Distill: Uncovering the True Intent behind Human-Robot Communication

As robots become increasingly integrated into everyday environments, intuitive communication paradigms such as natural language and end-user programming have become indispensable for specifying autonomous robot behavior. However, these mechanisms are ineffective at fully capturing user intent: natural language is imprecise and ambiguous, whereas end-user programming can be overly specific. As a result, understanding what users truly mean when they interact with robots remains a central challenge for human-AI communication systems. To address this issue, we propose the Distill approach for human-robot communication interfaces. Given a task specification provided by the user, Distill (1) removes unnecessary steps; (2) generalizes the meaning behind individual steps; and (3) relaxes ordering constraints between steps. We implemented Distill on a web interface and, through a crowdsourcing study, demonstrated its ability to elicit and refine user intent from initial task specifications.

cs.RO

Robust Sequential Experimental Design for A/B Testing

Experimental design has emerged as a powerful approach for improving the sample efficiency of A/B testing, yet existing designs rely critically on correctly specified models. We study robust sequential experimental design under model misspecification and develop a unified framework that covers both contextual bandit and dynamic settings. Theoretically, we prove that our design bounds the worst-case mean squared error of the estimated treatment effect. Empirically, we demonstrate the effectiveness of the proposed approach using synthetic and real-world datasets from a leading technology company.

stat.ML

Learning Perturbations to Extrapolate Your LLM

Recent advancements in large language models demonstrate that injecting perturbations can substantially enhance extrapolation performance. However, current approaches often rely on discrete perturbations with fixed designs, which limits their flexibility. In this work, we propose a framework where token prefixes are perturbed by a learnable transformation of a continuous latent vector within an embedding space. To overcome the challenge of an intractable marginal likelihood, we derive unbiased estimating equations for model parameters and optimize them via stochastic gradient descent. We establish the statistical properties of the resulting estimator in over-parameterized regimes. Empirical evaluations on both synthetic and real-world datasets demonstrate that our proposal yields significant gains in out-of-domain settings over a range of state-of-the-art baseline methods.

stat.ML

Can We Distinguish the Source Region Location of Filament/Prominence Eruptions from the Sun-as-a-star H$α$ Spectrum?

Solar filament/prominence eruptions can significantly perturb geospace when originating from favorable source locations and directions. While stellar analogs have been recently reported, the disk locations and magnetic environments of their source regions remain spatially unresolved on other stars. To bridge this gap, we investigate the typical Sun-as-a-star H$α$ temporal spectral characteristics of solar filament/prominence eruptions with different source region locations (on-disk vs. limb, active region vs. quiet-Sun region). It is revealed that limb eruptions are characterized by blueshifted/redshifted emission caused by the bright off-limb erupting structures, whereas on-disk eruptions may show blueshifted absorptions due to the dark erupting filaments. Among the limb eruptions, front-side limb eruptions usually display line center emission before the blueshifted/redshifted emission, while far-side limb eruptions show the opposite sequence. Moreover, the magnetic environment at source also shapes the spectral characteristics. On-disk filament eruptions from active region exhibit much more intense flare-ribbon-dominated line center emission features compared with those from quiet-Sun region. Limb active region eruptions often show single-wing emissions, whereas large-scale quiet-Sun region (quiescent) prominence eruptions frequently display expansion-induced emission in both wings followed by line center absorption due to the disappearance of bright prominence. These distinct Sun-as-a-star H$α$ spectral characteristics, dependent on eruption location, provide a diagnostic basis for inferring source regions of stellar filament/prominence eruptions from spatially unresolved H$α$ spectra.

astro-ph.SR

Magnetic Evolution of Highly-Sheared Region in Active Region 13842 Producing Large X9.0 Flare

Shearing motion and magnetic flux cancellation around the polarity inversion line (PIL) play significant roles in the build-up of free magnetic energy and magnetic flux rope (MFR) in source region of major solar flares. Here we investigate the magnetic evolution of a highly-sheared PIL in active region (AR) 13842, hosting the largest X9.0 flare of Solar Cycle 25. Since 2024 September 29, a positive-polarity pore persistently drifted northward along the western side of the AR's main negative-polarity sunspot. The main sunspot remained stationary until negative-polarity patches successively emerged to its east and approached. Rear-ended by these same-polarity patches, the sunspot then began moving westward toward the opposite-polarity pore around October 1, forming a collisional PIL. Meanwhile, on the PIL's other side, the pore was also rear-ended by same-polarity patches sequentially emerging behind it, accelerating the shearing motion around the PIL, where frequent flux cancellations were also observed. Synchronous rapid accumulation of free magnetic energy and formation of MFR were then observed in the PIL, where multiple major flares successively occurred within two days. Before these large flares, the area and total free energy of the high-free-energy-density PIL region gradually decreased in the photosphere, which could be caused by the initial ascent of MFR before eruption and serve as a precursor of solar eruptions. These results suggest that persistent flux emergences with cross separation directions facilitates rapid formation of collisional shearing PIL and frequent flux cancellations, leading to repeated MFR formations and multiple large flares in a relatively short time.

astro-ph.SR

Partially-shared Imaging Regression on Integrating Heterogeneous Brain-Cognition Associations across Alzheimer's Diagnoses

Alzheimer's Disease Neuroimaging Initiative (ADNI) diagnostic groups present strong heterogeneous associations among demographic, imaging, and cognitive data. We propose a novel PArtially-shared Imaging Regression (PAIR) model to represent imaging coefficients as weighted combinations of smooth spatial components. A Total Variation penalty is applied to enforce spatial smoothness, and a Selective Integration penalty is introduced to adaptively learn partial-sharing structures across groups. Theoretically, we establish minimax-optimal error bounds that dynamically adapt to varying sharing paradigms. Numerically, PAIR achieves predictive accuracy comparable to advanced deep learning models while providing superior interpretability. Applied to ADNI data, PAIR reveals substantial heterogeneity in brain-cognition pathways between cognitively normal (CN) and cognitively impaired (CI) groups, with hippocampal imaging contributing minimally in the CN group but substantially in the CI group, particularly in the CA1, CA3, and presubiculum subfields.

stat.ME

Solar Energetic Particle Events and Associated Type II Radio Bursts from Different Source Regions

Large solar energetic particle (SEP) events are thought to originate from the shocks driven by fast coronal mass ejections (CMEs) and thus generally accompanied by type II radio bursts. However, a significant proportion of type II radio bursts is not accompanied by SEP events. To study the relationship between SEPs and type II radio bursts and the associated physical mechanisms, we statistically analyze 43 SEP halo-CMEs and 131 non-SEP halo-CMEs observed from 2010 to 2024, and check the related properties of type II radio bursts and solar source region. We find nearly all SEP events and approximately two-thirds of non-SEP events are accompanied by type II radio bursts. Type II radio bursts associated with SEP events usually have longer duration and lower ending frequencies. The starting frequency exhibits a clear source region dependence, being highest for ''single active region (AR)'', intermediate for ''multiple ARs'', and lowest for ''outside of ARs''. Furthermore, the spectra of both protons and electrons exhibit a similar softening trend in the three types of source regions. Joint analysis of spectra and type II radio bursts reveals that the proton spectra index has a good anti-correlation with the starting frequency of the type II radio bursts. Our statistical results have important implications for the mechanisms behind SEP acceleration

astro-ph.SR

YASA: Scalable Multi-Language Taint Analysis on the Unified AST at Ant Group

Modern enterprises increasingly adopt diverse technology stacks with various programming languages, posing significant challenges for static application security testing (SAST). Existing taint analysis tools are predominantly designed for single languages, requiring substantial engineering effort that scales with language diversity. While multi-language tools like CodeQL, Joern, and WALA attempt to address these challenges, they face limitations in intermediate representation design, analysis precision, and extensibility, which make them difficult to scale effectively for large-scale industrial applications at Ant Group. To bridge this gap, we present YASA (Yet Another Static Analyzer), a unified multi-language static taint analysis framework designed for industrial-scale deployment. Specifically, YASA introduces the Unified Abstract Syntax Tree (UAST) that provides a unified abstraction for compatibility across diverse programming languages. Building on the UAST, YASA performs point-to analysis and taint propagation, leveraging a unified semantic model to manage language-agnostic constructs, while incorporating language-specific semantic models to handle other unique language features. When compared to 6 single- and 2 multi-language static analyzers on an industry-standard benchmark, YASA consistently outperformed all baselines across Java, JavaScript, Python, and Go. In real-world deployment within Ant Group, YASA analyzed over 100 million lines of code across 7.3K internal applications. It identified 314 previously unknown taint paths, with 92 of them confirmed as 0-day vulnerabilities. All vulnerabilities were responsibly reported, with 76 already patched by internal development teams, demonstrating YASA's practical effectiveness for securing large-scale industrial software systems.

cs.SE

Riemannian Motion Generation: A Unified Framework for Human Motion Representation and Generation via Riemannian Flow Matching

Human motion generation is often learned in Euclidean spaces, although valid motions follow structured non-Euclidean geometry. We present Riemannian Motion Generation (RMG), a unified framework that represents motion on a product manifold and learns dynamics via Riemannian flow matching. RMG factorizes motion into several manifold factors, yielding a scale-free representation with intrinsic normalization, and uses geodesic interpolation, tangent-space supervision, and manifold-preserving ODE integration for training and sampling. On HumanML3D, RMG achieves state-of-the-art FID in the HumanML3D format (0.043) and ranks first on all reported metrics under the MotionStreamer format. On MotionMillion, it also surpasses strong baselines (FID 5.6, R@1 0.86). Ablations show that the compact $\mathscr{T}+\mathscr{R}$ (translation + rotations) representation is the most stable and effective, highlighting geometry-aware modeling as a practical and scalable route to high-fidelity motion generation.

cs.CV