Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 613 records · Page 34Linked to original sources

SMAT: Simple and Efficient Merge-Aware Training

Model merging integrates the capabilities of multiple experts without joint retraining, but standard expert training optimizes task loss alone and does not guarantee good performance after merging. Merge-aware training (MAT) aims to improve merged performance, but existing methods do not fully account for common merging operations and add training cost. We observe that, from an expert's perspective, common merging methods can be described by three operations: Scale reweights its own update, Mask removes selected coordinates, and Perturb adds updates from other experts. Based on this view, we introduce SMAT (Simple MAT), which jointly optimizes expert loss and expected loss at simulated merged parameters generated by sampling scaling coefficients, masks, and additive noise. We further introduce periodic scheduling, kernel fusion, and parameter storage switching to make SMAT efficient, with one forward and one backward pass per step. Across four language and vision-language backbones, SMAT improves the mean score across five merging methods by 1.07-2.16 points over the strongest baseline for each backbone, with less than 2% training-time overhead over standard fine-tuning.

cs.LG↗

MA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement Learning

World models improve sample efficiency by training policies on imagined trajectories, but their usefulness depends on learning representations that capture the information needed for future control. We study whether self-supervised joint-embedding prediction (JEPA) can provide this learning signal for multi-agent reinforcement learning. We introduce MA-JEPA, a stochastic world model that replaces observation reconstruction with prediction of target representations, enabling model-based multi-agent reinforcement learning with centralized training and decentralized execution. A categorical latent state and a causal Transformer are trained with posterior and action-conditioned dynamics prediction objectives and are then used for actor-critic learning from latent imagination. A training-only joint predictor conditions on all agents' local states and actions to predict each agent's next local observation embedding. These predictions are passed through the same local posterior used during real interaction with a centralized critic that is used only for value learning, with execution remaining decentralized. Our experiments show that this architecture performs strongly on SMAC, matching or exceeding the strongest reported comparator mean win rate on four of eight evaluated maps.

cs.LG↗

Stochastic gravitational-wave background from self-interacting superradiant clouds

Gravitational-wave (GW) observations offer a powerful probe of new fundamental fields. One well-motivated source is black hole (BH)--boson cloud systems, in which an ultralight scalar field forms a cloud around a rotating BH via superradiance and emits long-lived nearly monochromatic GWs. In this work, we compute the stochastic GW background (SGWB) from such systems, extending previous work by including scalar self-interactions, and discuss its detectability with next-generation ground-based GW detectors. We find that self-interactions can suppress the SGWB and thereby relax existing LIGO-Virgo-KAGRA constraints inferred from null searches. Namely, the LIGO detectors at design sensitivity are insensitive to a SGWB produced by scalar fields with a decay constant $f_\mathrm{s} \lesssim 3\times 10^{17}$ GeV, independently of the scalar field mass. Looking ahead, we show that a moderately self-interacting cloud can still produce a detectable SGWB with next-generation detectors. Under conservative assumptions, for a decay constant $f_\mathrm{s} = 10^{17}$ GeV, the Einstein Telescope (ET) will be sensitive to scalar masses in the range $\sim[10^{-13.0},10^{-11.8}]$ eV, while Cosmic Explorer (CE) will be sensitive to scalar masses in the range $\sim[10^{-13.2},10^{-11.7}]$ eV. We also find that the minimum decay constants that still yield a detectable SGWB for ET and CE are $f_\mathrm{s}\sim 6\times10^{16}\,$GeV and $\sim 3\times10^{16}\,$GeV, respectively.

gr-qc↗

Opera: A Verbal Critic Framework for Long-horizon Coding Agents

Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is delivered. We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved. Opera decides when to review through periodic and event-driven triggers, diagnoses issues with typed operators, audits feedback against visible evidence before delivery, and tracks the agent's subsequent actions to distinguish mere compliance from actual resolution. As a test-time critic, Opera improves the resolve rate of non-critic agents by up to 12.4, 15.0, and 8.9 percentage points on Terminal-Bench 2.1, a SWE-Bench Pro subset, and DeepSWE v1.1, respectively, across four policy models, and achieves the highest mean resolve rate among competitive critic baselines on all three benchmarks, and also improves policy models when the policy critiques itself. Beyond inference, Opera-guided rollouts provide approximately on-policy training data: fine-tuning Qwen3.5-9B on them improves its resolve rate on held-out SWE-Bench Pro repositories by 10.2 percentage points without a critic at inference time, matching fine-tuning on rollouts from a stronger model, while preserving its performance when switching harness, i.e., from Openhands to Terminus-2, which the latter substantially degrades. Our code is available at: https://github.com/dongyuanjushi/Opera.

cs.CL↗

Efficient Reasoning via Constrained Optimization in Latent Space

Large Reasoning Models (LRMs) have shown remarkable reasoning capabilities, yet they still suffer from overthinking, generating redundant reasoning steps which incur substantial token consumption. Existing methods, such as suppressing reflective keywords or forcing shorter reasoning lengths, attempt to mitigate this issue but inevitably truncate necessary steps and induce underthinking, thereby compromising performance. To address this dilemma, we investigate the latent representations and observe that efficient reasoning steps naturally cluster into a concentrated region in latent space, while those deviating from this region tend to produce verbose sequences. To leverage this, we keep reasoning focused within this region via a quadratic program which projects deviating hidden states back into the region. Then we propose a novel training-free framework to achieve efficient reasoning that reduces token generation costs without sacrificing performance. Extensive experiments conducted on four models ranging from 1.5B to 14B, and across six benchmarks in math reasoning, coding, and scientific QA, validate the effectiveness of our method, up to a 12.1\% improvement in accuracy while reducing generated tokens by 11.8\% to 52.8\%. Codes are available at \href{https://github.com/hzn18/Opt4Reasoning}{https://github.com/hzn18/Opt4Reasoning}.

cs.AI↗

WorldWeave: Growing Persistent Geometric Worlds for Video Generation

Despite rapid progress, world models still lack explicit, persistent structural memory, making it difficult to preserve consistent world structure during continual scene expansion and cross-view revisits. To address this limitation, we present WorldWeave, a world generation framework that decouples world-state maintenance from visual rendering. Specifically, WorldWeave combines continual elevation-map generation with agent-guided scene organization and stitching to build an expandable explicit 3D world that incrementally extends structural memory while preserving existing structure. First, its terrain module uses diffusion-based image outpainting to generate continuous metric elevation maps under neighborhood conditioning and boundary constraints. Next, an agent integrates user intent, terrain evidence, and cross-region connectivity constraints to construct scenes through hierarchical semantic planning, deterministic geometry compilation, and local revision. Finally, during visual generation, planned camera trajectories query world geometry through a read-only interface, producing depth sequences that guide video synthesis without writing the generated results back into the world state. As a result, structural memory remains independent of short-window video generation, enabling continual expansion without predefined map boundaries and providing a consistent geometric basis for observations across trajectories and repeated visits.

cs.CV↗

Kaplansky decompositions of Polish modules

Let $R$ be a countable ring. Given an $R$-module $A$, we call a decomposition $A = \bigoplus_{i \in I} N_i$ a Kaplansky decomposition if each $N_i$ is countable. We characterize the uncountable Polish $R$-modules that admit a Kaplansky decomposition: they are exactly the modules of the form $B \oplus M^ω$, where $B$ and $M$ are countable and $M$ is $Σ$-algebraically compact. The countable summand $B$ may moreover be taken to be an elementary submodule satisfying a closure condition, which makes $M^ω$ unique up to isomorphism, and hence an invariant of $A$. We use this to characterize the countable rings admitting a free uncountable Polish $R$-module, generalizing results of Shelah and Solecki. This class of rings has a purely ring-theoretic description: it consists exactly of the countable left perfect and right coherent rings, i.e. the rings identified by Chase's theorem on products of projective modules. We observe that these are also exactly the countable $F$-rings, i.e. those countable rings $R$ for which $R^ω$ is free. Finally, we give a ring-theoretic characterization of the countable rings $R$ for which there exists an uncountable projective Polish $R$-module.

math.LO↗

FlowState: Execution State as Memory for Long-Horizon LLM Agents

Long-horizon tasks require LLM agents to continually draw on information from earlier interactions. However, retaining the full history increases context costs, while compressing it risks losing details needed later, and the relevance of historical information often becomes apparent as the task progresses. To address these challenges, we propose FlowState, which treats execution state as memory that can be retained and revisited across requests, unifying current decision-making with the reuse of historical information. FlowState preserves semantically typed state nodes, their relations, and references to raw tool observations, separating persistent retention from on-demand access. Within a single execution loop, Incremental State Update (ISU) maintains the current state based on new inputs and feedback, while Progressive State Access (PSA) progressively reveals historical states and supporting evidence as needed during reasoning. Together, these mechanisms enable agents to reassess prior decisions in light of new information and guide subsequent actions. Compared with a full-context baseline using the same DeepSeek-V4-Flash model, FlowState improves the average success rate on MemoryArena and the average pass rate on $τ^3$-Bench by 4.55 and 13.95 percentage points, respectively, while reducing total token consumption by 43.2% and 40.6%. These results demonstrate the performance and efficiency advantages of FlowState on long-horizon tasks.

cs.AI↗

Distances Between von Neumann Subalgebras: Spin Models, Commuting Squares, and Free Group Factors

We investigate the relative position of von Neumann subalgebras through the Mashood--Taylor ($\mathrm{d}_{\mathrm{MT}}$) and Kadison--Kastler ($\mathrm{d}_{\mathrm{KK}}$) distances and the interior angle. For each $n\in\mathbb{N}$, we show that the hyperfinite $\mathrm{II}_1$-factor $\mathscr{R}$ contains an uncountable family of pairwise distinct regular $n\times n$ spin model subfactors. In particular, this yields a continuous family $(\mathscr{R}_{\mathsf{H}_α})_{α\in[0,π)}$ of $2\times2$ spin model subfactors which, equipped with $\mathrm{d}_{\mathrm{MT}}$, is homeomorphic to the circle group $\mathbb{T}$ and admits a natural topological group structure. We further establish \[ \mathrm{d}_{\mathrm{MT}}(\mathscr{R}_{\mathsf{H}_α},\mathscr{R}_{\mathsf{H}_β}) =\vert\sin(α-β)\vert. \] We prove that two such spin model subfactors form a commuting square over their intersection if and only if they are maximally distant. More generally, commuting squares of $\mathrm{II}_1$-factors, under natural index conditions, force maximal distance, yielding $\mathrm{d}_{\mathrm{KK}}=1=\mathrm{d}_{\mathrm{MT}}$. We also show that diffuse subalgebras which are orthogonal in the sense of Popa are maximally distant. In contrast, no two members of the above spin model family are Popa-orthogonal. Nevertheless, maximal distance between two such subfactors implies that their interior angle over the intersection is $π/2$, highlighting the distinction between angle orthogonality and Popa orthogonality. Finally, in $L(\mathbb{F}_2)=L(\langle a,b\rangle)$, for $u\in L(\langle b\rangle)$, we obtain \[ \mathrm{d}_{\mathrm{MT}}(L(\langle a\rangle),uL(\langle a\rangle)u^*) =\sqrt{1-\vertτ(u)\vert^4}, \] and construct maximally distant masas not arising from subgroups of $\mathbb{F}_2$.

math.OA↗

Dynamical dark energy from quantum gravity condensates

We investigate dynamical dark energy emerging from quantum gravity within the framework of group field theory (GFT) condensate cosmology. We generalise previous models of quantum gravity induced cosmic acceleration by considering a broader parameter space, including an additional scalar matter field besides the relational clock, and accounting simultaneously for the phase dynamics and the contribution of multiple condensate modes. Working in a mean-field approximation with coherent peaked states and restricting to Hermitian GFT interactions, we derive the corresponding effective cosmological dynamics and analyse its late-time behaviour. We show that the system approaches an asymptotic de Sitter regime, dynamically dominated by a single condensate mode, while the approach to this regime can exhibit a non-trivial dynamical dark-energy evolution. In particular, both the condensate phase and subdominant modes can generate deviations from the asymptotic equation of state $w=-1$, including phantom behaviour and, under suitable conditions, a crossing of the phantom divide. The additional scalar matter field enters directly into these conditions, providing a link between the matter content and the quantum gravity induced dark-energy dynamics. These results extend the robustness and generality of the GFT mechanism for dynamical dark energy and provide a basis for confronting the underlying quantum gravity dynamics with cosmological observations.

gr-qc↗

Robust identification of drive-by sensing ride-hailing market with Points of Interest monitoring under uncertain ride demand

Taxi-based mobile sensing has emerged as a cost-efficient paradigm for large-scale urban environmental monitoring. In practice, both passive and active sensing strategies are adopted by ride-hailing platforms. Passive sensing, conducted during passenger-serving trips, is constrained by stochastic and spatially imbalanced ride demand, leading to limited and uneven coverage. Active sensing, executed by vacant taxis, provides greater control over sensing operations but incurs additional operational costs, and is thus typically treated as a supplementary strategy. To address the inefficiencies of the conventional "passive-first, active-second" paradigm, which may delay the monitoring of critical Points of Interest (POIs) and increase system-wide costs, we propose a Distributionally Robust Optimization (DRO)-based framework for active mobile sensing. First, we develop an enhanced A*-based drive-by sensing routing policy that integrates vehicle-task matching while capturing both global routing efficiency and local sensing opportunities. Second, we formulate the active sensing problem as a DRO model that explicitly accounts for uncertainty in ride demand through an ambiguity set of probability distributions, enabling robust decision-making for vacant taxi routing and vehicle-task matching. The proposed framework is evaluated on both static and dynamic mobile sensing settings using real-world data. Computational results demonstrate that our approach achieves better performance in terms of sensing coverage, operational cost, and robustness compared to benchmark strategies, highlighting the value of integrating distributional robustness into taxi-based sensing operations.

math.OC↗

Density functional perturbation theory of meta-generalized gradient approximations using algorithmic differentiation

Density functional perturbation theory (DFPT) is an established framework for the computation of derivatives in plane-wave density functional theory (DFT). We present an implementation of DFPT for exchange-correlation (XC) functionals $E_\mathrm{xc}(ρ,τ)$ that incorporate an explicit dependence on both the density $ρ$ and the kinetic energy density $τ$. This covers the popular class of semilocal meta-generalized gradient approximations (meta-GGAs) as well as broader nonlocal parametrizations. We sidestep the derivation of cumbersome XC second energy derivative expressions by recasting these derivatives as a Jacobian-vector product of the XC potentials, which we evaluate with algorithmic differentiation (AD) techniques. Integration with our previously developed AD-DFPT framework [N. F. Schmitz et al., npj Comput. Mater. 12, 6 (2026)] provides access to derivatives of arbitrary ground state quantities with respect to arbitrary perturbations. We employ AD-DFPT to compute a range of response properties for ZnO and BaTiO3, and find that the recent r2SCAN01 meta-GGA functional generally outperforms LDA and PBE. Finally, we showcase the optimization of a neural-network meta-GGA to self-consistently reproduce hybrid-DFT reference densities of bulk silicon, using AD-DFPT gradients. Overall, these results establish AD-DFPT as a versatile route for computing DFT derivatives at the meta-GGA level, be they common response properties or the unusual derivatives required for the gradient-based training of novel XC functionals.

cond-mat.mtrl-sci↗

From Reconnaissance to Response: Quantitative Risk Parameterization and Game Theoretic Containment in Modern Enterprise Attack

Modern Security Operations Centers struggle with delayed manual incident response, enabling adversaries to advance through the Cyber Kill Chain during early stage reconnaissance. While classical game theoretic defense models optimize strategic resource allocation, they rely on static utility matrices that fail to adapt to dynamic telemetry. This paper presents an integrated, metrics driven decision engine that bridges quantitative risk parameterization and continuous automated response time. Common Vulnerability Scoring Systems exploitability parameters are mapped to attacker success probabilities and evaluate defender log distributions via Factor Analysis of Information Risk Monte Carlo simulations. Real time SIEM logs streams are modeled as Poisson process arrival rates, dynamically updating defender posterior threat belief through sequential Bayesian filtering. A closed form threshold is derived by framing the interaction as a dynamic Bayesian Stackelberg game, where the expected unmitigated risk exceeds proactive containment cost. Parameterized against empirical data from the 2023 MGM Resorts and Caesars Entertainment cyber incident, simulation results demonstrate that the engine suppresses transient background noise while triggering automated SOAR network isolation within seconds of adversarial probing. Multi parameter sensitivity analysis confirms that the decision boundary dynamically adjusts to live perimeter vulnerability, offering a control theoretic foundation for sub minute automated threat containment.

cs.GT↗

Leveraging secondary-structure information for accurate nucleic acid structure prediction with OFoldNA

Recent advances in biomolecular structure prediction have enabled accurate modelling of increasingly complex molecular systems. However, nucleic acid structure prediction remains challenging because of conformational flexibility and the limited availability of high-quality 3D structural data. Secondary structure (SS) provides a more readily available layer of structural information that captures base-pairing relationships and folding topology. Here we present OFoldNA, an all-atom diffusion model that incorporates SS information into nucleic acid folding and protein--nucleic acid co-folding. Without external SS information, OFoldNA achieved leading performance on FoldBench for both nucleic acid monomer folding and protein--nucleic acid co-folding, with particularly strong performance on DNA monomers and protein--DNA interfaces involving longer nucleic acid chains. When accurate base-pairing information was provided, OFoldNA-SS2TS further improved both folding and co-folding accuracy, while partial SS information also yielded consistent gains. The same auxiliary branch can also be used for RNA SS prediction as OFoldNA-SS, which achieved the best out-of-distribution performance on CHANRG. Together, these results show that intermediate structural information such as nucleic acid SS can be leveraged to improve all-atom 3D modelling, providing a general direction for incorporating complementary structural modalities into molecular structure prediction and design.

q-bio.BM↗

Every countable group embeds in a group of type $\mathrm{FP}_n$

For every integer $n\ge2$, we prove that every countable group embeds in a group of type $\mathrm{FP}_n$. We also construct a group of type $\mathrm{F}_n$ containing a copy of every recursively presented group. Consequently, a finitely generated group embeds in a group of type $\mathrm{F}_n$ if and only if it is recursively presented. This answers questions of Fournier-Facio and Zaremsky, and confirms a suggestion of Gromov. An appendix combines the construction with the controlled-filling argument in a subsequently released OpenAI preprint to prove that every countable group embeds in a two-generator group of type $\mathrm{FP}_\infty$.

math.GR↗

CompOrca: Corpus-Scale Compliance Labelling of Instruction-Tuning Data

Studying how fine-tuning shapes refusal and noncompliance behaviour requires knowing which training examples refuse or otherwise fail to fulfil the request. Existing annotations cover evaluation sets, which are far smaller than training corpora. We present CompOrca, compliance labels for all 4,233,923 examples of the OpenOrca corpus. Every example was classified as compliant or noncompliant by five passes of an open-weight LLM judge (LongCat-2.0, 1.6T parameters). The corpus is released as unanimous compliance (94.75%), unanimous noncompliance (1.28%), and nonunanimous rows (3.97%), with the raw vote counts. A single pass flags 2.7-3.2% of the corpus as noncompliant, while only 1.28% is flagged by all five, so the most ambiguous rows can be filtered out. Against 450 human-annotated examples (150 annotated twice; human-human $κ=0.93$), the unanimous compliance and noncompliance labels are 97.3% and 86.7% precise. The noncompliance label is a high-precision subset of the corpus's noncompliance. Published refusal-detection methods recall between 0.4% and 94.1% of human-labelled noncompliance. We release the full corpus with its per-row labels and vote counts at https://huggingface.co/datasets/cemiu/CompOrca

cs.CL↗

Target-Aware Sequential Inference: Pooled versus Stratified Anytime-Valid Designs

Sequential studies with heterogeneous strata may define the estimand under a target mixture that differs from the sampling distribution. We organize the design problem by guarantee class: one pooled confidence sequence for a single prespecified target, shared pooled inference for a prespecified target set, and stratum-resolved inference when local guarantees must remain available. For target-set pooled inference we derive minimax, gap-aware, and range-capped shared proposals and an adaptive tracking rule. A specialized augmented off-policy score removes between-stratum mean heterogeneity asymptotically, recovering Neyman allocation for efficient single-target pooled inference. Robust directional certification over a convex target set is a classical intersection-union problem and does not require an M-fold vertex split. We derive a Gaussian information lower bound and an exact finite-support characteristic information that coincides with the bounded nonparametric KL-inf limit, and state sufficient conditions for first-order attainment of the Gaussian constant. Within the local-CS architecture, regularly varying boundaries imply a general allocation law with the 2/3 exponent for root-n variance-adaptive widths. Simulations and a PromptEval replay illustrate the resulting guarantee-efficiency frontier.

stat.ME↗

In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks

We study robotic in-context learning (ICL), an emerging paradigm that enables robots to infer and execute tasks from visual demonstrations. Despite its growing promise, the problem itself remains under-defined: a visual demonstration simultaneously conveys action trajectories, object semantics, manipulation affordances, spatial relations, and task goals, making it unclear what information the robot is actually expected to follow. In this work, we first provide a clear problem definition of robot ICL that explicitly defines its learning target and resolves this fundamental prompt ambiguity. Building on this definition, we develop a minimalist and reproducible ICL framework (SimpleICL) with a visual prompt encoder and a low-cost data collection protocol. Without massive pre-training or specialized data infrastructure, our framework achieves strong performance in both simulation and real-world environments. Extensive experiments further reveal several key properties of robot ICL, including semantic discrimination and task-relevant disentanglement. We will fully open-source our data and training pipeline to facilitate systematic and reproducible research on robot ICL. The project page can be found at https://simpleicl.github.io/simpleicl.

cs.RO↗