Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

SE-GoS: Self-Evolving Graph-of-Skills for Skill Library at Scale

LLM agents use large libraries of reusable skills. At thousands of skill entries, retrieval becomes the bottleneck. Graph-of-Skills (GoS) retrieves dependency-aware bundles from a typed skill graph, and SkillDAG shows that such a graph can accumulate execution-backed structure online. Neither asks whether execution traces can be distilled into a better retrieval graph that generalizes to unseen tasks. We present \textbf{Self-Evolving Graph-of-Skills (SE-GoS)}, which treats the retrieval graph as an index rather than a learned representation: the graph is maintained from execution traces while the retrieval pipeline, the skill library, and the model stay fixed. SE-GoS applies three updates: (1) \textbf{topology}, which induces relations from execution evidence and retracts an avoid edge only after repeated successful co-use; (2) \textbf{edge-weight}, which softly attenuates unsupported semantic edges and reinforces incoming edges to used skills; and (3) \textbf{node-description}, which updates retrieval-facing descriptions stored on graph nodes ranked too low. On SkillsBench, one evolution round lifts average reward from 52.4\% to 59.4\%, above full-library loading, vector retrieval, static GoS, and SkillDAG, and this ordering repeats on all three backbones. Retrieval over the evolved graph spends about two-thirds of the input tokens that loading the full library costs. Repeating the round does not help. The same graph improves a held-out split it never saw from 52.9\% to 58.3\%, so what it accumulates transfers rather than memorizes traces. Skill graphs can therefore be improved from execution experience without model training, retrieval-algorithm changes, skill-content modifications, or a model judging which skills are related.

cs.AI↗

Dynamics of Meaning: Towards the Evaluation of Diachronic Semantic Change in Sinhala

Tracking semantic change in low-resource languages across extensive historical timelines presents significant challenges due to data scarcity and the limitations of static embedding alignments. This study investigates the diachronic evolution of the Sinhala language from the 13th to the 20th century using a multi-stage computational framework. We first align century-specific Word2Vec and FastText embeddings using Similarity Matrix Based Alignment (SMA) and Orthogonal Procrustes (OP) techniques, finding that OP alignment provides more stable neighbourhood tracking for identifying temporal similarity dips. To move beyond aggregate measures, we introduce a Bidirectional Semantic Impact Pruning approach using contextualised embeddings from a fine-tuned Llama-3.1-8B. By applying Leave-One-Out (LOO) diagnostics, we attempt to isolate influential sentences to distinguish between systemic semantic shifts and transient polysemic expansion. Our results show that semantic drift in the fine-tuned Llama-3.1-8B is not evenly distributed across all usages. Instead, a significant part of the change is driven by a smaller set of high-impact contextual instances, rather than gradual and uniform change across all occurrences. This work provides a preliminary framework for low-resource Sinhala diachronic analysis, highlighting the trade-offs between model sensitivity and data availability.

cs.CL↗

On the cost of reaching an attainable state for the boundary-controlled 1-D heat equation

We consider the heat equation in an interval of the real line, with null initial condition, and Dirichlet boundary controls. We study the cost of reaching an attainable state at some time $T>0$. Precisely, given a state in the interval that we know is reachable by suitable boundary controls, we estimate the size of the smallest controls that reach it, in terms of an appropriate norm of this state. Our estimate is optimal in some sense as the time horizon goes to $0$.

math.OC↗

The derived length of subgroups of the Cremona group

We prove that the maximal derived length of a solvable subgroup of the plane Cremona group over $\mathbb C$ is equal to $5$. The main difficulty is the case of subgroups containing loxodromic elements. To treat this case, we prove that the maximal derived length of a finite solvable subgroup is $4$ and establish a centraliser theorem: the centraliser of a finite solvable subgroup of derived length $4$ contains no loxodromic element. The proof of the centraliser theorem relies on Blanc's classification of maximal algebraic subgroups and then uses equivariant Sarkisov theory, orbit estimates, superrigidity, and fixed curves of genus at least two.

math.AG↗

Impossibility of One-Way One-Round Quantum 4-Coloring via Matrix-Space Stability

We show that one-way one-round quantum-LOCAL algorithms cannot $4$-color directed cycles with high probability. This is the first lower bound in the high-probability quantum LOCAL setting that goes beyond the non-signaling and bounded-dependence models, exploiting the structure of distributed quantum algorithms. Our proof establishes a bidirectional connection between distributed quantum computing and extremal combinatorics. We obtain our lower bound by proving a Mantel-type stability theorem for weighted matrix spaces.

quant-ph↗

TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that model navigation as a 2D path planning problem, humanoid traversal in cluttered environments requires continuous geometry-aware whole-body adaptation, including coordinated arm placement, torso adjustment, and gait modulation for collision-free movement through complex 3D spaces. We introduce TANGO, the first whole-body vision-language navigation framework for language-conditioned humanoid traversal in cluttered environments. Given a natural-language instruction and egocentric RGB observations, TANGO directly predicts 29-DoF joint-space actions for downstream whole-body control. We train TANGO entirely in simulation by synthesizing diverse collision-free traversal behaviors via global path planning, kinematic whole-body motion generation, obstacle-aware motion editing, and RL-based tracking. This pipeline provides dynamically feasible action supervision for learning language-conditioned whole-body policies. In extensive simulation experiments, TANGO demonstrates state-of-the-art performance in vision-language navigation, while outperforming strong modular baselines in navigating challenging scenes requiring obstacle negotiation. Lastly, we deploy TANGO zero-shot on a Unitree G1 humanoid robot, and observe robust language-guided traversal in cluttered real-world scenes without training on any real-world navigation data.

cs.RO↗

Tensor-Train Weak SINDy: Identifying High-Dimensional Nonlinear Dynamics

Weak Sparse Identification of Nonlinear Dynamics (WSINDy) provides a noise-robust approach for learning dynamical systems from data without requiring numerical differentiation. However, for high-dimensional systems, tensor-product libraries of candidate functions grow exponentially with the state dimension, making standard WSINDy expensive in both computation and memory. The Multidimensional Approximation of Nonlinear Dynamics (MANDy) addresses this scaling through a tensor-train (TT) representation of the candidate library, but does not provide a mechanism for sparse model selection. Here, we combine these approaches to develop TT-WSINDy, which performs the weak-form transformation, regression, and sparsification in TT format. We show that the TT formulation recovers the corresponding WSINDy regression problem and derive polynomial time and memory complexity bounds for the tensor-train sparsification procedure. Numerical experiments demonstrate robustness to measurement noise and computational savings for high-dimensional systems.

cs.LG↗

ACT, WAIT, or EXPERIMENT: A Causal Governance Framework for Retail Price Optimization Under Abstentions

This paper presents a causal decision-making framework for estimating price elasticity in retail channels, a process typically confounded by promotions, competitor movements, and market frictions. Rather than forcing a calculation when data is ambiguous, the system introduces decision abstention (\textsc{wait}) as an active diagnostic tool rather than an estimation failure. Combining Double Machine Learning and conformal prediction, the tool evaluates whether reliable conditions exist to adjust prices or if pausing the decision is preferable. When the system abstains, it exhaustively classifies the reason for the pause, identifying which products require designed pricing experiments or whether aggregating data to the brand level restores usable estimates. Tested on controlled synthetic data, the model shows that this operational discipline drastically reduces estimation error (lowering RMSE from 0.571 to 0.159) and offers a practical, secure alternative to blind estimation in thin-data retail environments.

econ.EM↗

Structural Sign Herdability in Temporal Networks: A Sufficient Condition via $π_p$-Graphs

In this letter, we study the herdability of temporally switching directed networks. A temporal network is modeled as a switched system with a fixed switching sequence, which imposes more restrictive herdability conditions than those of conventional switched systems. By exploiting the relationship between temporal walks and the entries of the controllability matrix, we derive sufficient conditions for herdability. We further show that the magnitude of edge weights influences the sign pattern of the controllability matrix, thereby affecting herdability. Consequently, herdability in temporal networks depends not only on the network topology and switching durations, but also on the magnitude of the edge weights. Motivated by this observation, we establish equivalent graph-theoretic conditions for structural sign ($\mathcal{SS}$) herdability in temporal networks. In particular, we introduce the union multigraph of temporal subsystems and propose the notion of a $π$-graph. We show that the existence of a $π_p$-graph, which is a temporally evolving $π$-graph, is sufficient to guarantee $\mathcal{SS}$ herdability. Illustrative examples are provided to demonstrate the proposed results.

eess.SY↗

CHOIR: heterogeneity-aware conformal prediction for crash injury severity across driver safety strata

Transportation agencies increasingly predict crash-injury severity with statistical and machine-learning models, but these models do not state how often their output contains the recorded injury level or for which groups of drivers it fails, a gap that matters most for motorcyclists and unrestrained drivers. This study develops and evaluates a certification layer that gives any fitted severity model a finite-sample, distribution-free coverage guarantee within prespecified safety strata. The layer, CHOIR (Conformal Heterogeneity-aware Ordinal Inference with Risk control), combines groupwise and weighted conformal prediction with conformal risk control to return contiguous KABCO intervals, and adds a declared sensitivity analysis for medically assessed injury and bounds on fatal omission. It is evaluated on 4.04 million Texas crashes from 2017-2023, one sampled driver per crash, with seven base models from the ordered logit to a tabular foundation model, and on held-out counties and later years. Under one pooled threshold every model reaches 0.90 coverage overall but covers motorcyclists or unrestrained drivers at 0.868 or lower, and class-balanced gradient boosting covers unrestrained drivers at only 0.374. Calibration within four safety strata places all 28 model-by-stratum estimates between 0.898 and 0.907, at the cost of sets spanning 3.6 to 4.6 of five categories for these groups, and the certified ordered logit is within 0.03 categories of the narrowest model. Injury-model coverage should therefore be certified within safety groups rather than on average, calibration rather than model complexity determines validity, and a statewide threshold should not be applied to small rural counties without local calibration data.

stat.ML↗

Asymptotic $q,t$-Fuss--Catalan numbers for type $B$

Let $W=W(B_n)$ act diagonally on $\mathfrak{h}\oplus\mathfrak{h}^*$, let $S=\mathbb{C}[\mathfrak{h}\oplus\mathfrak{h}^*]$, let $J\subset S$ be the ideal generated by the $W$-alternating polynomials and $\mathfrak{m}_S$ is the maximal ideal of the origin. For sufficiently large $m$ we compute $q,t$-Fuss-Catalan polynomial $Cat^{(m)}(B_n;q,t):=Hilb(\frac{J^m}{\mathfrak{m}_S J^m})_{det-part}$ and imply $Cat^{(m)}(B_n;1,1)=\binom{n(m+1)}{n}$. For proofs, we work with the $Γ$-equivariant Hilbert scheme $Y_n=nΓ$-$Hilb(\mathbb{C}^2)$, $Γ=μ_2$ and Haiman-type Koszul complex that defines the punctual locus of $Y_n$. Our formula for $Cat^{(m)}(B_n;q,t)$ is derived from a localization computaion for the Haiman-type Koszul complex.

math.CO↗

Strategic Interview Positioning under Uncertain Self-Rank

We study interview positioning in the secretary problem when an applicant is uncertain about their own absolute rank. The employer uses a given rejection threshold, and the applicant's self-rank distribution is truncated geometric. We show that the applicant need only compare the first eligible position with the final position and that, for finite $n$, a unique critical value exists precisely when $r(r+1)>n-1$. If $r_n$ is the employer's rejection threshold and $r_n/n\toα\in(0,1)$ as $n\to\infty$, then \[ \lim_{n\to\infty}q_c(n,r_n) = q_c(α) = \frac{1-α}{1-α+α^2}. \] Thus uncertainty about absolute rank produces an explicit asymptotic phase transition in optimal interview positioning.

math.PR↗

Scaling Clinical Judgment to Evaluate Medical AI

Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus, prior studies typically rely on small physician panels, often from a single institution or specialty, which both limits the scientific questions investigated and makes it unclear whether findings would be reproduced with a different set of evaluators. To more rigorously and scalably study clinical reasoning in AI models, here we introduce PrecepTron, an LLM fine-tuned for physician-level evaluation of open-ended responses. PrecepTron was trained using low-rank adaptation (LoRA) of a 32-billion-parameter model on a small number of physician examples. We also release GRAND-ROUNDS, a new large-scale physician-annotated benchmark of 9,217 scored responses from 160 clinicians across seven studies. We show that frontier LLMs in typical "LLM-as-a-judge" approaches often disagree with physicians and with each other, but fine-tuning PrecepTron on a small number of cases enables physician-level consistent scoring across tasks. We use PrecepTron to reproduce headline findings from five influential studies assessing LLMs for clinical care in JAMA, Science, and Nature Medicine without new human grading. Using PrecepTron, we then pose new questions about how LLMs reason in medicine that would have been infeasible with human grading alone, including measuring the diagnostic accuracy of frontier LLMs when clinical cases are provided piecemeal, even token by token. Together, PrecepTron and GRAND-ROUNDS provide a foundation for reproducible, large-scale study of how LLMs reason in medicine. All code, data, and labels are made freely available for researchers.

cs.AI↗

Inverse Problem of Alchemical Resource Theory: Replication and Universal Simulation Single Out Imaginarity and Parity Asymmetry

Quantum resource theories usually begin with a prescribed free structure and ask what tasks become possible when a resource is supplied. We study the inverse problem of alchemical resource theories, which we define as resource theories admitting a resource that can both replicate itself exactly and universally simulate quantum instruments. Under natural consistency assumptions and branchwise complete freeness, we show that for multi-qubit systems only two nontrivial theories survive, up to a common local unitary change of basis: parity asymmetry and imaginarity. We further show that exact universality exhibits an all-or-nothing trade-off: any nonmaximal resource state can exactly simulate only free unitaries. These results demonstrate that prescribed operational capabilities can strongly constrain the underlying resource structure and reveal a fundamental trade-off between computational capability and the precision required for physical implementation.

quant-ph↗

A lattice family with kissing numbers $τ(\mathcal{L}_n) \ge e^{2 \sqrt{n}}$

For all prime powers $q\geq5$, we construct lattices $\mathcal{L}_q\subseteq\mathbb{Z}^q$ with kissing numbers \[ τ(\mathcal{L}_q)\geq \left(\frac{1}{2πe^2}+o(1)\right)\sqrt{q}\,e^{2\sqrt{q}}. \] The same asymptotic bound holds on a set of integer dimensions of natural density $1$, and in every sufficiently large integer dimension $n$ with an additional factor $e^{-\tfrac{1}{2}n^{1/40}}$. The construction is an extension of a previous construction by Bennett-Peikert based on Reed-Solomon codes.

math.CO↗

Windowed A-K-MDP

Markov decision processes (MDPs) are used to support decision-making in conservation of biodiversity, but policies, even over small state spaces, can be difficult to interpret for conservation managers. K-MDP methods address this problem by building simpler MDPs with at most K abstract states. We show that the previously proposed A-K-MDP algorithm that relies on selecting a discretisation divisor using binary search can skip better abstract states. To fix this issue, we propose Windowed A-K-MDP, an algorithm that generates every distinct feasible partition induced within a declared divisor window and evaluates candidates until reaching the ideal value loss (J = 0) or exhausting the family of candidates. Across 33 K-MDP instances, Windowed improved 25 and tied 8.

cs.AI↗

Near-thermal state and thermodynamic response of a QCD system in ultra-central heavy-ion collisions

We present a systematic framework to extract QCD thermodynamic properties at finite baryon density from ultra-central heavy-ion collisions, where the impact parameter is nearly zero and volume fluctuations are strongly suppressed. By mapping the measured multi-particle state of each event to an effective homogeneous fireball through the conservation of total energy, total entropy and net baryon number, a near-thermal state with variations is realized from event to event, which allows thermodynamic response relations to be tested. Using STAR Beam Energy Scan data on charged particle multiplicity and net proton fluctuations, we show that the thermalization condition is satisfied within uncertainties across collision energies, validating the interpretation of ultra-central events as small deviations from a thermally equilibrated state. We then generalize the response relations to finite baryon chemical potential, expressing the response coefficients in terms of the speed of sound and susceptibilities. Our results provide a baseline for using ultra-central collisions to probe the QCD equation of state and phase structure, including the search for the critical endpoint.

nucl-th↗

Spectral density of angular momentum transfer from a swift electron to a large spherical nanoparticle

Swift electrons in scanning transmission electron microscopy transfer both linear and angular momentum to nanoparticles, underlying electron-beam-driven nanoscale manipulation ("electron tweezers"). Prior theory relied either on the small-particle (dipolar) approximation, valid only well below experimentally relevant sizes, or on frequency-integrated multipolar calculations that leave the spectral structure of the interaction unresolved. Here we present a fully retarded, causal, multipole-converged electrodynamical methodology for the angular momentum transfer from a swift electron to an isolated spherical nanoparticle, based on a closed-surface Maxwell stress tensor formulation whose angular integrals reduce analytically to a small, material- and trajectory-independent set of irreducible integrals over associated Legendre functions. This lowers the cost of the double multipolar sum, enabling convergence up to l_max=51 for nanoparticles as large as a=50 nm, nearly four times the order of the largest previous calculation at this size and previously unreached for an optically complex, interband-dominated material. Applied to aluminum and gold nanoparticles up to a=50 nm, the method resolves the transfer spectral density across the full frequency domain, showing it is set by interference between the electron field and the field scattered by the nanoparticle, which dominates the scattered-scattered contribution at essentially every frequency; the electric contribution dominates at moderate speeds, but the magnetic contribution, of the same sign, grows steadily in relative weight with electron speed, from a few percent of the total at v = 0.5c to between a quarter and a half of it at v = 0.95c for trajectories passing close to the nanoparticle surface.

cond-mat.mes-hall↗