Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Constructing Disambiguated Knowledge Bases from Large Language Models at Scale

Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language models (LLMs), treating the model itself as the knowledge source. However, LLMs natively possess no representation of entities, leading to duplicate entries as well as conflations. We propose GPTKB 2.0, a methodology for constructing disambiguated KBs directly from LLMs. GPTKB 2.0 incorporates on-the-fly disambiguation of entities, relations and classes, and is meticulously designed to satisfy both scalability and disambiguation accuracy. We analyze the central design decisions and characterize the trade-offs between accuracy, scale, and cost. We execute GPTKB 2.0 at scale, obtaining a materialized KB containing over 1M disambiguated entities and 38.4M triples. This represents the first million-scale LLM-native KB with explicit internal canonicalization of entities, relations, and classes, a significant departure from prior Wikimedia-centric works. GPTKB 2.0 is available at https://gptkb.org/.

cs.CL↗

Trie Constraints and Hierarchy-Aware Semantic Alignment for HS Code Prediction with Small Language Models

Harmonized System (HS) code prediction (HSP) from commodity text is essential to international trade, and its importance continues to grow in port logistics. Recently, large language models (LLMs) have been actively investigated for this task, owing especially to their strong language-understanding capabilities. However, their high computational cost limits deployment in constrained environments such as container terminals. Small language models (SLMs) offer a practical alternative, but their smaller scale makes them prone to generating invalid HS codes and to overlooking the hierarchical semantics between commodity text and HS codes. To address these limitations, this study proposes TRIE-HSA, which combines trie-constrained token prediction with hierarchy-aware semantic alignment (HSA). This framework constrains the SLM to predict only valid digits under the HS taxonomy and aligns commodity text representations with the hierarchical structure of HS codes. In extensive experiments on data collected from an operational container terminal, TRIE-HSA improved average HS6 accuracy by 49.96% over zero-shot inference and exceeded the strongest task-specific benchmark by 11.94%. These results demonstrate that accurate and structurally valid HSP is achievable with fewer than 10 billion parameters. Therefore, TRIE-HSA offers a practical basis for deployment of HSP in port logistics operations that cannot support large scale LLMs.

cs.CE↗

Defensive Pessimism: Growth, Culture, and the Survival of Evidence

A common monetary standard can expand credit while weakening the alternative practices needed to recover from its failure and identify its cause. I study a growing credit economy in which safeguarded real settlements supply recovery capacity, matched observations for a reform decision, and apprenticeships in an alternative practice. The efficient reserve grows in absolute size but becomes a vanishing share of the economy. Protection based on a fixed population share therefore eventually fails, even along a path with expanding capacity and evidence. Voluntary cooperation provides only bounded readiness, while cheaper drills can preserve recovery without producing evidence. Apprenticeship and expected political survival reinforce one another, allowing exposed and protected continuations. I characterize protection indexed to the required service and the least-cost funding of replacement if the original institution is cancelled. The resulting arrangement combines verified procurement, training support when needed, and safe real collateral held by an outside custodian.

econ.TH↗

Optimal Training-Time Scaling in Gradual Adaptation

In gradual adaptation, how should the training time on each task change as the number of intermediate tasks increases? We study this question for overparameterized linear regression tasks that change smoothly and share a zero-loss solution. With $N$ tasks and training time $s_N$ on each, the final learning progress converges to a continuum curve when $Ns_N\toτ$. The limiting progress is $Θ(τ)$ for small $τ$ and $Θ(τ^{-1})$ for large $τ$, so both very short and very long training produce little progress. It follows that optimal per-task training times scale as $s_N^\star=Θ(N^{-1})$, equivalently $Ns_N^\star=Θ(1)$. Experiments on gradually rotated MNIST and a natural Yearbook time shift are consistent with less per-task training as the path is divided more finely.

cs.LG↗

omni-macos: On-Device Omni-Modal Search on Apple Silicon

We present omni-macos, a search engine that embeds text, code, documents, images, audio and video into one representation space and runs its encoder, index and store on the Mac that already holds the files, so no indexed file, no typed query and no vector ever leaves the machine. It keeps a background indexer and an interactive search box inside one memory budget the user sets: it embeds and stores each distinct chunk once, re-encodes only the chunks an edit changes, hands the GPU smaller units while the user is typing, answers queries from a one-bit replica of the index with exact rescoring, and propagates that budget to the allocators that draw on unified memory. We measure on five Macs spanning an eightfold range of accelerator width and a thirty-twofold range of memory, each indexing the files it already holds.

cs.IR↗

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding

Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows. Since the appropriate frame budget varies with the downstream LMM, reasoning demands, and latency constraints, a practical selector should serve multiple budgets. However, existing methods typically optimize an isolated frame subset for each predefined budget: when the budget changes, previously selected evidence may be replaced rather than progressively augmented. A fixed-weight ranking allows prefix reuse across budgets but applies the same weighting at every position, overlooking the distinct roles of early and later ranks. We formulate long-video frame selection as a Matryoshka ranking problem: constructing a single priority sequence whose small prefixes concentrate query-conditioned evidence, while progressively larger prefixes preserve this evidence and add broader temporal context. Efficiently constructing such a ranking is itself challenging, as densely sampling long videos and evaluating frame-query relevance incurs substantial overhead. We therefore introduce Matryoshka Evidence-to-Context (MEC) Frame Selection, a training-free framework that builds a reusable sparse video index, discovers candidates through sparse probing and local zooming, and greedily constructs a position-adaptive ranking - early positions emphasize evidence; later positions progressively favor temporal coverage while preserving visual diversity. A single ranking can thus be truncated to any target budget without rerunning the selector. Across four benchmarks and six frame budgets, MEC improves average accuracy over uniform sampling by 3.77 points, matches strong state-of-the-art selectors, and reduces end-to-end selection latency by 47.37-51.19%.

cs.CV↗

Schreier Sets of Intervals, Super-Schreier Sets, and Catalan Numbers

A finite nonempty set $F\subset\mathbb{N}$ is Schreier if $\min F\ge |F|$. First, we prove a linear recurrence relation and compute initial counts for Schreier sets consisting of intervals. Two intervals of integers are separated if their union is not an interval. If $\mathcal J_{k,n}$ is the collection of Schreier sets that are the union of exactly $k$ separated intervals, then the sequence $(|\mathcal{J}_{k,n}|)_{n=1}^\infty$ satisfies the characteristic polynomial $p_k(x) = (x-1)^{2k+1}(x+1)^k$. Furthermore, we introduce the new concept of $k$-super-Schreier sets and let $\mathcal{S}_{k,n}$ denote the collection of $k$-super-Schreier sets whose maximum is $n$. We show that the sequence $(|\mathcal{S}_{k,n}|)_{n=1}^\infty$ satisfies a Fibonacci-type recurrence with a remainder term expressible as a polynomial of $n$.

math.CO↗

EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning

Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on external simulators that are difficult to ground. We introduce EnvACE, an agentic reinforcement learning method that replaces external environment interaction during training with world rehearsal. The policy alternates between acting and rehearsal: it first generates a tool call, then plays the role of the environment to produce the response induced by that action, and conditions subsequent decisions on the rehearsed response. Both roles are jointly optimized end-to-end using task-success rewards. Through world rehearsal, the policy internalizes the relationship between actions and their environment responses in its parameters, yielding an agent world model that directly supports decision making. Across BFCL-v4, tau^2-Bench, VitaBench, and FinMCP-Bench, EnvACE achieves strong and transferable performance, outperforming environment-scaling baselines in the overall evaluation. Controlled studies further show that world rehearsal consistently improves policy learning across model scales. At test time, the internalized world model enables private rehearsal before committed execution, yielding further gains under a moderate rehearsal budget without additional external interaction. Our findings establish world rehearsal as a new path toward scaling LLM agent training beyond the constraints of external environments. Our code is publicly available at https://github.com/Within-yao/EnvACE.

cs.AI↗

Exponential logical-error reduction in quantum memories via optimal syndrome-measurement timing

Syndrome-measurement timing is usually treated as a fixed clock cycle of a quantum error-correcting code. For quantum memories, however, the inter-round interval is itself an optimizable control parameter: measuring too rarely allows idling errors to accumulate, whereas measuring too often introduces measurement-induced faults. We propose a phenomenological logical-noise model for this trade-off and analytically show that the optimal syndrome-measurement interval is inversely proportional to the code distance and that this produces an exponential reduction of logical-error rates in the distance relative to constant-interval schedules. Furthermore, for time-dependent idling noise, we develop an adaptive timing strategy based on the measured syndrome activity that outperforms every fixed-interval protocol, with largest gains for short but strong noise bursts. Simulations of rotated surface-code memories with matching decoding validate the phenomenological model, the distance-dependent optimum, and the adaptive-strategy improvement. Moreover, with the experimental noise parameters reported by Google in Nature 638 (2025), our model predicts reductions in logical-error rates per unit time of up to $40\%$.

quant-ph↗

Hypothesis Testing with Conditional Queries: Learnability and the Value of Interaction

Model evaluations may fix all tests before observing any responses or select later tests using earlier responses. We study this choice in a conditional-query model on a finite outcome space $\mathcal{X}$ with $|\mathcal{X}|=N$. We first ask which pairs of distribution classes can be reliably distinguished. We then ask how many additional queries are required to match an adaptive tester when all queried events must be fixed in advance. We show that learnability holds if and only if the two classes have positive separation in their pairwise conditional probabilities. When this separation is zero, the optimal worst-case error is exactly $1/2$ at every finite query budget. For any $T$-query adaptive policy and any $ρ\in (0,1)$, we construct a randomized non-adaptive procedure using $O(N^2(T + \log(1/ρ)))$ pair queries chosen before any response is observed. Its simulated transcript is within $ρ$ in total variation of the adaptive transcript, uniformly over all distributions in the model. We also construct a matching family with constant adaptive query complexity and $Ω_\varepsilon(N^2)$ non-adaptive query complexity. Consequently, the worst-case fixed-error adaptivity gap is $Θ_\varepsilon(N^2)$. Thus interaction can reduce the required number of tests by a quadratic factor, but the apparent exponential branching of an interactive evaluation does not yield an exponential query advantage.

cs.LG↗

SPT-3G D1: Foreground-Robust Lensing Templates for Primordial Gravitational Wave Searches

Gravitational lensing of the cosmic microwave background (CMB) generates B-mode polarization that acts as a source of contamination to searches for B modes generated by primordial gravitational waves (PGWs). The strongest constraint on PGW B modes is already significantly limited by lensing B modes, as shown in the most recent BICEP result. In this work, we present CMB lensing B-mode templates constructed using SPT-3G and Planck data, which characterize the lensing B modes and can be used to improve PGW B-mode searches. We use SPT-3G data from the 2019 and 2020 observing seasons for the E modes and the CMB-reconstructed lensing potential, and a cosmic infrared background (CIB) map from Planck as an external lensing tracer. To test for extragalactic foreground biases in the lensing template, we consider CMB lensing reconstruction variants with different levels of foreground immunity: the standard and profile-hardened global minimum variance (GMV) quadratic estimators, and a polarization-only quadratic estimator. We validate the template construction using Gaussian simulations and Agora simulations with realistic non-Gaussian foregrounds. From simulations, we find that foreground-induced biases are strongly suppressed for the template constructed with the profile-hardened GMV + CIB tracer, with residual bias below 10% of the statistical uncertainty on the template power spectrum. Data difference tests on this template similarly show no evidence for significant foreground contamination. This foreground-immune lensing template achieves delensed residual BB power of $A_{\rm lens}^{\rm res} \simeq 0.48$ averaged over $20 \leq \ell \leq 200$, the highest delensing efficiency lensing template to date. These results demonstrate and validate a method to construct foreground-robust lensing templates which will be used in upcoming delensed PGW B-mode analyses of BICEP data.

astro-ph.CO↗

Dimension-Free Polylogarithmic Quantum Shadow Tomography

Shadow Tomography is a fundamental problem in quantum information theory. Given multiple copies of an unknown $d$-dimensional quantum state $ρ$ and a known collection of observables $E_1,\ldots,E_M$, the goal is to estimate all expectation values $\{\text{Tr}(ρE_i)\}_{i=1}^M$ to additive accuracy $\varepsilon$ with probability at least $1-δ$. An elusive open question from the seminal shadow tomography work of Aaronson is whether this task admits a dimension-independent sample complexity with only polylogarithmic dependence on $M$, as suggested by the best-known lower bounds. In this work, we propose two different quantum protocols for shadow tomography with the best sample complexity \[ O\left( \frac{\log(M)\log(M/δ)}{\varepsilon^2} \right), \] which is polylogarithmic in the number of observables and independent of the dimension of the unknown state, thereby answering Aaronson's original question while also providing an exponential improvement in the prior best dimension independent sample complexity of shadow tomography from Sinha (STOC 2025) and, more recently, Chen, O'Donnell, Pelecanos, and Wright. Our approach first reduces the general shadow tomography problem to a finite-ensemble estimation problem via a minimax argument. We then develop an observable-independent protocol that repeatedly applies the pretty-good measurement while updating the prior distribution over the finite ensemble according to the measurement outcomes. A tail analysis of the resulting estimation error yields simultaneous accuracy guarantees for all observables and a cubic-logarithmic upper bound. We also introduce a refined recovery-label measurement for the same finite ensemble, which yields the bound in our main theorem.

quant-ph↗

Fast LapSum: Exact Differentiable Top-$k$ at Million Scale

Selecting the top-$k$ elements is a fundamental operation for inducing sparsity in large-scale models and optimization problems, enabling robust expert activation, token routing or attention pruning. However, hard top-$k$ is non-differentiable, while existing differentiable alternatives become increasingly expensive as the number of coordinates grows. We introduce Fast LapSum, a scalable solver for the LapSum soft top-$k$ formulation that preserves an exact selection mass of $k$, while supporting end-to-end differentiation. In Fast LapSum, we reduce sorting cost using probabilistic bracketing, which restricts sorting to a narrow band of scores around the threshold using a binomial order-statistic from kernel-noised samples. A certification pass upgrades the probabilistic localization to a verified one at the cost of one additional linear pass, with a full-sort fallback that covers the worst-case scenario. Our certified GPU implementation processes $10^6$, $10^7$, and $10^8$ scores in median times of $0.92$, $1.39$, and $7.24$ ms, respectively, making exact-budget soft top-$k$ practical within million-scale optimization loops. We demonstrate this capability in two applications: megapixel sparse adversarial examples with a small fraction of initial image pixels, where Fast LapSum achieves an order-of-magnitude speedup over state-of-the-art methods, and 3D Gaussian splatting. In the latter, we use Fast LapSum to reduce the number of Gaussians produced by Adaptive Density Control to a substantially smaller number while retaining nearly the same rendering quality and massively reducing computation. These results demonstrate that exact-budget differentiable top-$k$ can be incorporated into practical million-scale optimization pipelines.

cs.AI↗

Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models

Mixture-of-Experts (MoE) models decouple parameter count from per-token compute, but deployment still requires hosting every expert in memory. Recent theory shows that experts whose router weights change least during fine-tuning can be pruned with provable accuracy preservation, yet the guarantee assumes full fine-tuning. We show that the signal can be elicited through a brief parameter-efficient adaptation. We fine-tune with a lightweight adapter, rank experts by the induced router change, and prune the least-changed experts in one shot. On Mixtral-8$\times$7B-Instruct, router-only LoRA trains 0.002% of parameters and retains 27.54% MMLU-Pro accuracy with half the experts removed, against roughly 16% for magnitude and random pruning. Signal quality improves monotonically with adapter size, reaching 28.76%, and declines as adaptation spreads beyond the router. Under their shared budget, IA3 reaches 28.04% while Houlsby reaches 25.39%. The criterion transfers to Qwen1.5-MoE fine-tuned for mathematical reasoning, retaining 49.7% mean accuracy over eleven benchmarks with half the experts removed. Structural pruning reduces memory by 49% and per-token latency by 37%. Lightweight router sensitivity therefore makes provably motivated, task-conditioned expert pruning practical at scale.

cs.LG↗

Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways

Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial intelligence. Currently, mechanistic interpretability studies on multilingual safety are largely confined to local components, such as isolated neurons. However, this static and fragmented perspective overlooks the synergy among components and fails to elucidate how safety signals dynamically propagate within the model to drive safety decisions ultimately. In this work, we move beyond isolated neurons to identify and target the cross-layer functional pathways formed during safety signal propagation, thereby uncovering the mechanisms driving the cross-lingual safety gap. Specifically, we first identify monolingual safety pathways and validate their impact on refusing harmful requests. Subsequent cross-lingual analyses reveal a sparse subset of cross-lingual shared safety pathways, confirming that this intersection acts as the internal bridge transferring safety capabilities from high-resource (HR) languages to non-high-resource (NHR) languages. Building on these mechanistic findings, we propose a pathways-targeted alignment method based on the cross-lingual shared safety pathways. Experimental results show that updating only a small fraction of pathway parameters significantly improves safety in NHR languages while largely preserving the model's general capabilities.

cs.AI↗

poly_rust: a domain-specific language for polytopal methods with a Rust interpreter

Background: Over the last two decades, considerable effort in the applied mathematics community has been devoted to the development of polytopal methods, i.e., numerical methods for partial differential equations that support meshes more general than those used by traditional finite element methods. Despite their flexibility and attractive mathematical properties, polytopal methods have a steep implementation curve and, to date, no easily accessible open-source software framework has been available for their systematic implementation. Methods: In this article, we introduce poly_rust, a new software library implementing a domain-specific language (DSL) specifically designed for polytopal methods. poly_rust provides high-level constructs for expressing a wide range of such methods using a syntax that closely mirrors their mathematical formulation, thereby substantially lowering the implementation barrier for both newcomers and experienced users. Results: The library facilitates the rapid implementation, testing, and comparison of polytopal discretizations, making these methods more accessible to researchers, applied scientists, engineers, and students. In our vision, this will foster their dissemination beyond the numerical analysis community and encourage their adoption by practitioners in a broad range of application fields. It will also contribute to reproducibility of results, as DSL files are text files that can be easily made available, stored, and exchanged. Conclusions: poly_rust introduces, to our knowledge, the first DSL specifically tailored to polytopal methods. By combining mathematical expressiveness, ease of implementation, and an open-source framework, it provides a common platform for developing, sharing, reproducing, and comparing polytopal schemes, with the ambition of becoming a reference tool for the wider scientific community.

math.NA↗

Schwinger-Keldysh effective field theory of type-B Goldstone: near-diagonal geometry and Berry term

In this work, we formulate a finite temperature Schwinger-Keldysh effective field theory for type-B Goldstone modes based on geometric objects globally defined on the coset manifold. In particular, after taking the classical limit, we organize the required two time contour structure by a near-diagonal geometry in which the $r$-type field is interpreted as the physical Goldstone configuration and the $a$-type field is identified as corresponding tangent displacement. Within this geometric viewpoint, the Berry term essential for type-B Goldstones is obtained directly through the transgression of the Berry curvature. Moreover, we show that this exact Berry transgression is compatible with dynamical KMS condition. In order to construct the conservative, dissipative and noise sectors, we classify possible tensors globally defined on the coset manifold. Based on this framework, we discuss in detail two concrete model examples. The dispersion relations of the Goldstone modes and the associated two-point correlation functions are calculated in the presence of dissipation.

hep-th↗

Disproving the Petersen Coloring Conjecture: Theoretical Analysis and an Infinite Family of Counterexamples

In 1988, Jaeger conjectured that every bridgeless cubic graph $G$ admits a Petersen coloring; that is, a map $E(G) \to E(P)$ mapping any two adjacent edges of $G$ to two adjacent edges of the Petersen graph $P$. A positive resolution of Jaeger's conjecture would have immediately resolved several other famous and long-standing problems in graph theory. In July 2026, a 68-vertex counterexample was announced on X. Shortly afterwards, Putman independently presented two non-isomorphic 112-vertex counterexamples, relying solely on computer-assisted verification. In this paper, we present two counterexamples of order $52$, currently the smallest known, and provide a purely theoretical proof. In the second part, we construct an infinite family of cyclically $4$-edge-connected cubic graphs without a Petersen coloring for every even order at least $60$. Additionally, through computational verification, we show that any counterexample must have order at least $40$. Moreover, we show that our counterexamples provide a negative answer to other related problems. Finally, we conclude the paper by discussing key open problems and highlighting avenues for future work.

math.CO↗