Search arXiv⌕ Search

arXiv subjects

Yong Li

Publications and source records attributed to Yong Li.

At least 19 recordsLinked to original sources

Optimal rates of uniform convergence for weighted Birkhoff averages via almost all rotations

In this paper, we investigate weighted Birkhoff averages for toral translations associated with compactly supported weighting functions. By introducing several new analytical techniques, we establish optimal uniform convergence rates for almost all rotations and specific (or even all) initial points. Unlike the $\mathcal{O}(N^{-1})$ rate best achieved in classical ergodic theory, we show that these weighted averages exhibit polynomial or even exponential convergence. We establish the optimality of these convergence rates in multiple aspects, particularly concerning regularity indices across four distinct cases: finite differentiability, the $C^\infty$ class, logarithmic $C^\infty$ classes, and Gevrey classes. Our results demonstrate that the regularity of the observable essentially dictates the convergence rate; furthermore, we prove that no admissible choice of weighting function can, in general, overcome the lower bounds imposed by this regularity. In contrast to the generically slow convergence of standard time averages, this work provides an optimal and nearly complete characterization of rapid convergence for weighted Birkhoff averages.

math.DS↗

WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory and In-Context Learning

World Action Models (WAMs) offer a promising paradigm for robotic manipulation by jointly modeling visual state transitions and robot actions. However, existing WAMs are constrained by limited temporal context, coarse episode-level language supervision, and predominantly text-only conditioning, which hinder task-progress tracking and fine-grained language-video-action grounding while limiting visual-context reasoning and cross-embodiment transfer. In this paper, we introduce WorldScape Policy 2.0, a controllable WAM with reasoning-augmented long short-term memory. Its causal short-term visual memory supplies recent observations as DiT prefill to preserve local interaction dynamics, while its long short-term event memory organizes historical VLM outputs into global-history, local-active, and event-boundary representations for progress-aware retrieval. The retrieved history augments perception and autoregressively generated planning tokens, yielding an implicit subgoal condition for autonomous planning; semantic forcing further transfers event-level instruction semantics into this latent planning pathway. To establish fine-grained multimodal controllability, we construct ManipEvent-5M, an event-grounded embodied pretraining dataset containing nearly 5 million event segments with aligned action trajectories, episode-level task instructions, segment-level subtask captions, goal images, and video demonstrations. These designs provide a unified interface for autonomous planning from high-level instructions and controllable execution from fine-grained text, goal-image, or video-context prompts. Experiments in both simulation and real-world platforms demonstrate superior capabilities in long-horizon autonomous planning, fine-grained instruction following and in-context adaptation.

cs.RO↗

WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models

Despite rapid progress in interactive world models (IWMs), short-horizon performance does not establish sustained action following, visual stability, physical plausibility, or memory. We introduce WorldRoamBench, an open-world benchmark for long-horizon stability across four dimensions, each with innovations: (i) Action: per-frame action metric bypassing cross-model semantic scale disparity and exposing failures hidden by trajectory; (ii) Vision: sliding-window drift metric capturing non-monotonic mid-sequence collapse missed by start-vs-end comparisons; (iii) Physics: evaluation of physical plausibility across mechanics, optics, and 3D consistency, gated by camera-motion and subject-tracking checks; (iv) Memory: a trajectory-aware protocol reducing confounding from action-following errors, evaluating scene memory via transition-localized 3D point-cloud reconstruction and subject memory via tracking-plus-VLM reasoning. The benchmark comprises 1000+ test cases across Nature, Urban, and Indoor scenes in first/third-person views with WASD 10-60 s continuous interaction. Evaluating 10+ open/closed-source models reveals none reliably satisfies all dimensions; even the best achieves only moderate scores. Advances on WorldRoamBench are steps toward IWMs that are stable, physically grounded, memory-faithful, and deployable in real-world applications.

cs.CV↗

Complete characterization of planar $C^1$ linearizability and optimal conjugacy regularity

We establish a complete characterization of local $C^1$ linearizability for planar diffeomorphisms. For every prescribed hyperbolic real Jordan form, we obtain necessary and sufficient integrability conditions on the modulus of continuity of the derivative for local $C^1$ linearizability. These criteria reduce to three canonical thresholds: the classical, polynomially weighted, and squared-logarithmically weighted Dini conditions. Each threshold is optimal: whenever the corresponding condition fails, we construct a diffeomorphism with the identical linear part that is not locally $C^1$ linearizable. For every non-hyperbolic linear part, explicit polynomial counterexamples show that smoothness alone cannot guarantee even local topological linearization. Beyond $C^1$ existence, we determine the optimal regularity for the derivative of the linearizing conjugacy in every hyperbolic case, with the optimality established by counterexamples.

math.DS↗

Libra: Taming Attention Workload Skew in Long-Context LLM Training with Bounded Sequence Pool

Long-context LLM training suffers from a load-balancing problem that sequence packing does not solve. Packing samples into fixed-token sequences balances memory and linear-cost operators, but the dominant attention cost scales with the sum of squared sequence lengths. Thus, equally sized packed sequences drawn from a long-tailed corpus can carry substantially different attention workloads, creating data-parallel stragglers and pipeline bubbles. Existing approaches either balance at the granularity of sequences or microbatches, where an outlier can dominate an assignment, or disaggregate attention over a global worker pool whose communication domain grows with the data-parallel (DP) degree. We present Libra, which operationalizes the law of large numbers (LLN) as a scaling principle for load balancing: the attention-balancing pool need not grow with the DP degree. Libra groups packed sequences and their CP groups into fixed-size sequence pools. As DP scales out, Libra adds pools rather than enlarging each one, bounding every attention exchange. Variance-Reduced Sequence Placement makes this effective for finite, long-tailed workloads by co-locating sequences with complementary attention workloads to reduce residual inter-pool skew. Within each pool, Tiled Attention Pooling dispatches sequence-head SH-Tiles across GPUs, while a pipelined runtime overlaps tile exchange with attention. Libra exposes a drop-in context-parallel attention operator and a pluggable data sampler, requiring no changes to model layers, optimizers, or pipeline schedules. On three production Qwen3 models (8B, 30B, 235B) and 256K- and 1M-token production workloads, Libra improves end-to-end training throughput over the strongest evaluated baseline (WLB-LLM) by 44% on average and up to 68% at 256 GPUs. Libra has run for hundreds of thousands of GPU-hours in production on jobs spanning 32K to 1M tokens.

cs.DC↗

Programmable Hong--Ou--Mandel interference in a giant-atom beam splitter

The Hong--Ou--Mandel (HOM) effect is a hallmark of two-photon quantum interference, in which two indistinguishable photons impinging on a balanced beam splitter bunch into the same output port. Here, we show that a giant atom (GA), coupled to two waveguides through two coupling points each, can function as a programmable HOM interferometer, enabling continuous control over single-photon beam splitting and two-photon interference. This programmability arises from coupling-phase differences in the GA, which tune its self-interference and directionality, and thereby its scattering response. To characterize the two-photon interference, we analyze the scattering of two Gaussian single-photon wave packets injected through different waveguides and evaluate the bunching and antibunching (coincidence) probabilities of the resulting four output ports. At the operating point where the GA acts as an effective 50:50 beam splitter for single photons, we observe a pronounced HOM dip as the relative input delay between the two wave packets is varied. Away from this point, adjusting the coupling phases continuously tunes the two-photon output statistics between bunching into the same output port and antibunching across distinct output ports. In addition, we explore an application to quantum parameter estimation, showing that small deviations of coupling phases can be estimated from the two-photon output statistics, with the achievable sensitivity quantified by the classical Fisher information associated with a binary coincidence measurement. Our giant-atom beam splitter thus provides a programmable platform for two-photon interference in waveguide quantum electrodynamics, with potential applications in quantum information processing, quantum communication, and quantum sensing.

quant-ph↗

SeetaPsych v1.0: An Open-source Computer Vision Toolkit for Behavior-based Psychological Measurement

Automated visual analysis opens new avenues for behavior--based psychological measurement. Nevertheless, existing technological modules are typically scattered across task specific systems with heterogeneous interfaces and disparate deployment requirements. In this work, we present SeetaPsych v1.0, an open source, unified and extensible computer vision toolkit designed to extract psychologically relevant signals from facial images and/or face based videos. The current release encompasses four major core modules aiming at behavior--based physiological perception: unified face based emotion analysis (simultaneous facial expression recognition, facial action unit detection, and valence--arousal estimation), camera based heart rate estimation, screen point--of--gaze estimation, and scene gaze following. A suite of auxiliary preprocessing modules for human centric visual analysis is also included, comprising face detection, facial landmark detection, and head detection. These functionalities are encapsulated within a modular Pipeline/Runner architecture that automatically resolves attribute dependencies, constructs computation graphs, and support intermediate result sharing among modules. SeetaPsych provides standardized Python APIs to facilitate reproducible, large scale analyses, alongside an interactive WebUI for rapid, code--free method evaluation. Overall, SeetaPsych offers an integrated and accessible visual measurement platform for research in psychology, behavioral science, human computer interaction, and related fields.

cs.CV↗

When Cognitive Graphs Meet LLMs: BDEI Cognitive Pathways for Panic Emotional Arousal Prediction

Predicting the timing of individual and collective panic emotional arousal before manifestation is essential for timely emergency intervention. Existing methods incorporate cognitive elements but none of them model emotion in the generative direction of the arousal process, leaving arousal timing undetermined. We argue that grounding prediction in appraisal emotion theory is necessary because it models this process explicitly in its natural generative direction, but three problems must be solved. (1) Appraisal theory posits that emotion arises from simultaneous evaluation across multiple threat dimensions, yet no prior work fuses these inputs into risk perception; (2) Existing models are trained in the opposite, behavior-bridged direction, recovering emotion merely as a post-hoc correlate of behavior; (3) Approaches that adopt LLMs as the primary decision-maker yet overlook the fragility and hallucination-proneness of their outputs. We introduce PanicCognitivePath (PCP) to address all three. A Psychological Safety Distance (PSD) model, grounded in psychological distance theory, maps four-domain signals (physical, social, cognitive, and informational) into a unified risk metric that gates entry to cognitive reasoning. An explicit Emotion node grounded in appraisal emotion theory is introduced into BDI, forming a novel Belief-Desire-Emotion-Intention (BDEI) pathway that couples threat appraisal directly to emotional arousal. Inverting the conventional LLM-as-decision-maker paradigm, PCP confines the LLM to parameter estimation for the Belief-to-Desire transition, restricting hallucinations to a single step and curbing their accumulation across steps. Experiments on Hurricane Sandy show PCP improves individual prediction accuracy by 10.68% over baselines, reduces peak count error to 7.07%.

cs.CL↗

Experimental observation of exceptional bound states in the continuum

We experimentally demonstrate second- and third-order exceptional bound states in the continuum (EP-BICs), formed by the merging of two and three symmetry-protected BICs at an exceptional point (EP). Our passive reciprocal acoustic platform consists of symmetry-protected BIC cavities coupled through an acoustic waveguide and enables independent control of intrinsic loss, radiative loss, and near-field coupling between the cavities. A nonuniformly distributed intrinsic loss provides the non-Hermiticity required for EP formation while preserving decoupling from the waveguide radiation channel. The EP-BICs are probed using both near- and far-field excitation. In the near field, local intracavity excitation directly accesses the symmetry-protected BICs and reveals their spectral evolution. For far-field excitation, we intentionally break the protecting geometrical symmetry, converting the BICs into quasi-BICs and making them accessible through transmission measurements. Extending the system to three cavities with graded intrinsic loss realizes a third-order EP-BIC and yields a larger spectral response to the implemented coupling perturbation. These results establish a passive reciprocal route to higher-order exceptional degeneracies with independent control over intrinsic loss, near-field coupling, and radiative access.

physics.optics↗

Persistence and long-time breakdown of most probable paths under time-dependent fractional noise with applications to KAM tori

We investigate the persistence of most probable paths through the Onsager--Machlup functional for multidimensional stochastic differential equations driven by fractional Brownian motion with time-dependent diffusion coefficients and Hurst parameter $H\in(1/4,1)$. Under suitable structural and variational conditions, deterministic trajectories remain most probable paths for sufficiently small noise in both the fixed-endpoint transition problem and the free-endpoint evolution problem, whereas sufficiently large noise destroys their local minimality. More generally, when exact persistence does not hold, global most probable paths converge to the corresponding trajectories of the noise-free system in both the uniform and Hölder topologies at the rate $O(ε)$. We further analyze the second variation along periodic deterministic trajectories over time intervals of length $NT$. For $H>1/2$, positive definiteness, and hence local minimality, is lost on sufficiently long intervals. For $H\in(1/4,1/2]$, long-time positive definiteness holds for the fixed-endpoint problem, but this conclusion does not directly extend to the free-endpoint setting. We also establish the persistence of KAM tori in nearly integrable Hamiltonian systems in the sense of most probable evolution paths. Finally, a two-dimensional numerical example illustrates the persistence of deterministic trajectories under small noise and their pronounced deviation under large noise.

math.PR↗

Emergence of polarization in networks of large language model agents

Rapid advances in large language models (LLMs) have not only empowered autonomous agents to generate social networks, communicate, and form shared and diverging opinions on political issues, but have also begun to play a growing role in shaping human political deliberation. Our understanding of their collective behaviours and underlying mechanisms remains incomplete, however, posing unexpected risks to human society. In this paper, we simulate networked systems involving thousands of LLM agents across different backbone models (GPT-3.5, GPT-4o, ChatGLM, Llama-3, and DeepSeek-V3), in which agents interact through LLM-guided conversations and update their opinions over time, resulting in the emergence of opinion polarization. We discover that these agents spontaneously develop their own social network with properties characteristic of human social networks, including homophilic clustering. The collective opinions of these LLM agents evolve in ways that exhibit behavioural patterns consistent with social phenomena and mechanisms widely discussed in studies of human behaviour. Overall, these behaviours, patterns, and emergent phenomena produced by LLM agents are consistent with real-world observations and established opinion-dynamics models. This consistency suggests that LLM agents can serve as a valuable synthetic testbed for exploring hypothetical intervention strategies in networked LLM-agent systems.

cs.SI↗

A globally defined polyconvex isotropic energy satisfying the true-stress-true-strain monotonicity condition (TSTS-M++)

Polyconvexity is a standard ingredient in the variational existence theory of finite elasticity, whereas true-stress-true-strain monotonicity (TSTS-M++) requires a positive incremental Cauchy-stress response. These two constitutive restrictions are independent, and Wollner, Holzapfel and Neff left open whether a compressible isotropic energy defined on the whole of $\mathrm{GL}^{+}(3)$ can satisfy both. We give an explicit affirmative answer. For every $μ>0$ and $k>0$, the stored-energy function $W_k(F)=\fracμ{2k}\bigl[\exp\bigl(k(\lVert F\rVert^2+3J^{-1}+J-7)\bigr)-1\bigr]$, $J=\det F$, is polyconvex and strictly rank-one convex. Its Cauchy-stress response satisfies TSTS-M++ globally if and only if $k\ge 1/(8\sqrt{3})$. In this regime every symmetric Cauchy stress corresponds to a unique positive-definite stretch, while the reference stretch is stress free with positive infinitesimal shear and bulk moduli. Stress bijectivity has the strictly smaller sharp threshold $k_{\mathrm{B}}\approx 0.00827233304$: at equality the stress map is a global homeomorphism with a nondifferentiable inverse, and above it the map is a global $C^{\infty}$ diffeomorphism. Thus, for $k_{\mathrm{B}}\le k<1/(8\sqrt{3})$, the map $V\mapstoσ(V)$ remains globally bijective while TSTS-M++ fails at finite strain. In the TSTS-M++ regime, every prescribed inner radius of a finite plane-strain annulus with a traction-free outer wall has a unique radial equilibrium, and its inner pressure increases smoothly and strictly from zero to infinity.

math.AP↗

BER-PEF: Unified Human Mobility Predictability Evaluation via Bayes Error Rate Estimation

Human mobility predictability concerns the best prediction performance attainable from a given target and input information, but its ground truth is not directly observable on real mobility data. We present BER-PEF, a Bayes-error-rate-based framework that converts BER estimation into mobility predictability estimation and provides a unified protocol for comparing estimators without observable ground truth. The framework maps symbolic sequences, numeric trajectories, contextual features, and learned representations into a common feature--label space, then evaluates estimator outputs along controlled perturbation curves against a shared predictability reference interval by measuring deviations below the interval, above the interval, and across the full interval. Experiments on Foursquare NYC and TKY, GeoLife, and T-Drive show that several BER-based estimators achieve lower reference discrepancy than existing predictability methods on symbolic sequences and numeric trajectories, while their estimates track changes in empirical prediction performance under perturbation. Additional analyses show that contextual inputs and multiple structured representations can be evaluated under the same protocol, and that aggregating evidence across multiple perturbation levels provides a more reliable basis for estimator selection than relying on a single unperturbed observation. BER-PEF therefore offers a unified and verifiable path for evaluating predictability estimators on heterogeneous mobility data when ground-truth predictability is unavailable.

cs.LG↗

AI agents reshape consensus formation in human groups

As large language model (LLM) agents shift from tools to participants in human groups, a fundamental question for collective behavior is how their growing presence reshapes consensus formation. Here we study mixed human-AI groups in a collaborative description game, in which shared conventions emerge through repeated rounds of random pairwise communication. Varying the proportions of LLM agents, we identify three distinct regimes of consensus formation: low agent proportions facilitate human-led consensus, intermediate proportions disrupt convergence, and high proportions restore strong consensus while shifting it toward agent-led conventions. Crucially, these regimes differ not only in the strength of convergence, but also in the semantic grounding and communicative form of the resulting consensus: human-led consensus is more concrete, holistic, and grounded in shared real-world analogies, whereas agent-led consensus is more abstract, less information-dense, and more geometrically segmented. Mechanistically, agent influence arises from a shared linguistic prior that places agents near one another in the expression space, combined with relatively stable expression choices across rounds; humans initially resist adopting expressions from partners perceived as AI but gradually yield to conformity pressure. These findings provide evidence that AI composition can shape the emergence, content, and perceived legitimacy of group norms, making agent proportion and transparency important design variables for human-AI systems.

cs.CL↗

Vitrification-Devitrification Enables Tunable Photonic and Gas Sorption Properties of Zeolitic Imidazolate Frameworks

Zeolitic imidazolate framework (ZIF) glasses represent an emerging family of melt-quenched glasses, which exhibit immense potential for applications in gas separation, energy storage, and optics. However, their intrinsic porosity remains elusive due to the inherent challenges in resolving their disordered atomistic structures. Here, we systematically investigate the porosity of ZIF-4 and ZIF-62 crystals and their corresponding glasses. CO2 sorption at 195 K enables quantitative assessment of microporosity in both crystalline and glassy states, allowing the accessible micropore volume of the ZIF glasses to be determined. Moreover, establishing a direct relationship between photonic properties and structural porosity in Zn-based ZIF glasses remains challenging. Here we demonstrate striking broadband blue-light emission from ZIF-4 glass annealed under optimized conditions. A pronounced red shift is observed when increasing the annealing temperature above the glass transition temperature. By correlating the evolution of photoluminescence with structural porosity, we reveal the interplay between the photonic and gas sorption properties of ZIF glasses. These findings provide new insights into the structure-property relationships of ZIF glasses and offer a pathway toward the rational design of multifunctional MOF glasses.

cond-mat.mtrl-sci↗

UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks

Large language model (LLM) agents have been widely applied in automating data science tasks. However, existing methods typically rely on a limited set of provided datasets, and they face challenges in data-intensive scenarios that require discovering and leveraging relevant information from large-scale and heterogeneous data repositories. Urban tasks are representative examples of such scenarios, as urban data are not only large-scale and multi-sourced, but also exhibit complex spatial, temporal, and semantic relationships. To address these challenges, we propose UrbanDS, a graph-guided LLM multi-agent system for data-intensive urban tasks. We first construct a unified dataset graph to organize reusable dataset skills and the relationships among datasets. Specifically, we develop a Data Profiling Agent that constructs a skill for each dataset. Moreover, a Relation Agent identifies relationships among datasets and integrates these relationships into the dataset graph. At runtime, a Planner Agent retrieves task-relevant datasets from the graph and generates execution plans. Multiple Execution Agents then perform data processing and analysis, while their execution progress and intermediate results are shared through a common memory. Finally, a Report Agent synthesizes the experimental logs into a report, which can be further refined based on user feedback. To systematically evaluate the capability of agents in handling data-intensive urban scenarios, we further construct UrbanDS-Bench, an urban data science benchmark covering representative data analysis and modeling tasks. Experiments on both general and urban benchmarks demonstrate that UrbanDS consistently outperforms existing data science agents on data-intensive tasks. Furthermore, UrbanDS has been deployed on the urban operations platform of Dongxihu District, Wuhan, demonstrating its effectiveness in real-world urban applications.

cs.AI↗

CateKV: On Sequential Consistency for Long-Context LLM Inference Acceleration

Large language models (LLMs) have demonstrated strong capabilities in handling long-context tasks, but processing such long contexts remains challenging due to the substantial memory requirements and inference latency. In this work, we discover that certain attention heads exhibit sequential consistency in their attention patterns, which can be persistently identified using a coefficient-of-variation-based algorithm. Inspired by this observation, we propose CateKV, a hybrid KV cache method that retains only critical token information for consistent heads, thereby reducing KV cache size and computational overhead, while preserving the majority of KV pairs in adaptive heads to ensure high accuracy. We show the unique characteristics of our algorithm and its extension with existing acceleration methods. Comprehensive evaluations on long-context benchmarks show that, while maintaining accuracy comparable to full attention, CateKV reduces memory usage by up to $2.72\times$ and accelerates decoding by $2.18\times$ in single-sample inputs, and boosts throughput by $3.96\times$ in batch scenarios.

cs.LG↗

CAER: Causal Action Effect Reweighting for World Model Training

World models are becoming core infrastructure for embodied intelligence, with action-conditioned video generation providing controllable predictions of how scenes evolve after agent interventions. Yet existing models are commonly trained with space-time-uniform mean squared error, allowing abundant background tokens to dominate the gradient while sparse interaction dynamics remain under-optimized; such uniform fitting rewards reconstructing appearance rather than learning how actions change the world. We introduce Causal Action Effect Reweighting (CAER), a general training paradigm that redistributes supervision toward the tokens whose predicted future is causally affected by the action. CAER contrasts the model's own predictions with and without action conditioning to localize these tokens online, then normalizes the resulting effect map into a weight that preserves the total coefficient mass and changes only where it is spent. This online signal requires no external annotations or offline preprocessing, avoids additional data-processing time, and scales naturally with model and dataset size. Experiments across heterogeneous action-conditioned world-model tasks show that CAER converges to better solutions than uniform MSE training, with consistent improvements in the physical consistency, controllability, and visual quality of generated videos.

cs.AI↗