Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 379 records · Page 21Linked to original sources

Decentralized Multi-Agent Systems with Shared Context

Multi-agent systems (MAS) can scale large language model agents on long-horizon tasks by running them in parallel, yet existing designs waste much of this parallelism in bubbles: agent time spent waiting on others or redoing a peer's work. These bubbles stem from how agents communicate. Independent agents share nothing and rediscover what their peers have already found; peer-communicating agents wait at synchronous rounds; and under centralized orchestration, the main agent blocks on its sub-agents while progress is relayed. We propose Decentralized Language Models (DeLM), a MAS framework on top of existing agent harnesses that squeezes out these bubbles by replacing the main agent with a shared context and a task queue. Agents asynchronously claim tasks, publish findings as soon as they are available, and build on or correct one another's progress, with every peer's status visible to all. On long-horizon tasks from Terminal-Bench 4.0 and DeepSWE v1.1, and on SWE-bench Verified, DeLM is both more accurate and faster than Codex, Claude Code, their native subagents, and AOrchestra in every setting, improving accuracy by up to 17.5 points over the strongest baseline and running up to 2.49x faster than the harness it builds on. On ProgramBench, where agents rebuild programs from scratch, DeLM makes faster progress than Claude Code and finishes a 120-minute budget up to 19.9 points higher in test pass rate. The code is available on our project website at https://yuzhenmao.github.io/DeLM/.

cs.MA↗

Derivatives of Tensor Products and Applications to Spaces $C(K^n,X)$

In this paper, we develop an abstract theory of derivatives for Banach spaces based on objects that we call \emph{bidual assignments}. This framework encompasses both the Semadeni derivative and the recently introduced Semadeni--Pełczyński derivative. More generally, suitable ideals of subsets of dual spaces give rise to a broad family of derivatives within this setting. We establish direct-sum and tensor-product formulas for these derivatives, showing that they behave naturally with respect to direct sums and injective tensor products. We then obtain explicit descriptions of derivatives associated with compact trees and study the corresponding questions for finite products of compact lines. As the main application, we establish classification results for spaces of the form $C(K^n,X)$. In particular, for uncountable ordinals $α$ and $β$, an integer $n\geq1$, and Banach spaces $X$ satisfying suitable rigidity assumptions, we prove that \[ C([0,α]^n,X)\sim C([0,β]^n,X) \quad\text{if and only if}\quad C([0,α])\sim C([0,β]). \] This extends Kislyakov's classification of the spaces $C([0,α])$ and its vector-valued extension due to Galego to finite powers of ordinal intervals.

math.FA↗

Thermal contact and direct detection in a broken $Z_4$ Majorana dark matter model

We consider a broken $Z_4$ Majorana dark matter model in which the fermion mass is generated by a real singlet scalar. For small scalar mixing, freeze-out is dominated by $χχ\to h_2h_2$, whereas momentum exchange with the Standard Model bath becomes inefficient. In this region, the usual assumption $T_χ=T$ is not always adequate. We therefore evolve the dark matter abundance and kinetic temperature simultaneously. Although $T_χ$ is still close to the bath temperature around chemical freeze-out, the later departure from kinetic equilibrium reduces the residual $p$-wave annihilation rate and leaves a larger relic abundance. In the sub-TeV region considered here, the difference between the two-moment and single-temperature calculations reaches $37.1\%$ at fixed parameters. Refitting the Yukawa coupling to reproduce $Ω_χh^2=0.12$ increases the spin-independent scattering cross section by up to $18.6\%$. A momentum-resolved Fokker--Planck calculation gives a similar shift, $17.7\%$, at a representative point. The LZ upper limit on the scalar mixing changes by less than $0.3\%$, since kinetic equilibrium is already well maintained in that part of parameter space. The lower mixing boundary obtained from the comparison of the two relic calculations therefore indicates where the common-temperature approximation loses accuracy; it is not an experimental lower bound. These results apply to the specified low-reheating broken-phase cosmology and to the region where threshold and strong nonrelativistic-potential effects remain small.

hep-ph↗

Deformable In-Hand Slip-Aware Tactile Sensor with Integrated Velocity Sensing, Force/Torque and Pressure Map Estimation

This paper introduces a novel tactile sensor for in-hand manipulation with slip-aware control that integrates velocity and force/torque sensing with pressure map estimation into a single device with a deformable contact pad. To the best of our knowledge, this is the first sensor to combine these sensing modalities within a single compliant structure. The sensor features a deformable contact surface and can robustly track both flat and curved surfaces across a wide range of diffuse surface materials. Its performance is evaluated through a comprehensive set of experiments that highlight both its capabilities and limitations. The sensor is designed for rapid and low-cost fabrication using a combination of standard PCB manufacturing and rapid prototyping techniques.

cs.RO↗

Quasi-Product States and Factor Types for the One-Dimensional Hard-Core Model

We study the von Neumann algebra generated by a quasi-product state associated with the one-dimensional hard-core lattice gas. If \(κ>0\) denotes the activity (fugacity), the transfer-matrix construction gives the stationary transition matrix \[ P=\begin{pmatrix}p&q\\ 1&0\end{pmatrix}, \qquad p+q=1, \qquad κ=\frac{q}{p^2}. \] The stationary Markov measure on the hard-core path space induces a faithful diagonal state on the corresponding path \(AF\)-algebra. We give a state-preserving identification of its GNS von Neumann algebra with the Feldman--Moore algebra of the measured tail equivalence relation. The relation is hyperfinite and ergodic, and its nonzero ratio set is \(κ^{\mathbb Z}\). Consequently, the generated algebra is the hyperfinite factor of type \(\mathrm{II}_1\) when \(κ=1\), and of type \(\mathrm{III}_λ\), where \(λ=\min\{κ,κ^{-1}\}\), when \(κ\ne1\). The point \(κ=1\) is not a thermodynamic phase transition: the pressure and particle density are analytic for all \(κ>0\). We also describe the precise support-corner relation with the GNS algebra of the diagonal Gibbs state on the quasi-local spin algebra, and we determine the centralizer and the flow of weights.

math-ph↗

Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning

When large language models (LLMs) fail to generalize or make content-sensitive errors in reasoning, it is often taken as evidence that LLMs are not truly reasoning, but rather performing a kind of pattern matching. The implication is that human behavior does not exhibit the same types of failures because human reasoning relies on principled and content-invariant world models. We test this assumption by first evaluating humans and LLMs on their ability to engage in common-sense reasoning about a variety of everyday situations. Our results reveal convergent patterns of reasoning across 46 LLMs and two cohorts of human participants. We then ask whether this behavioral convergence is due to LLMs having acquired content-invariant world models or a set of pattern-matching heuristics by characterizing the roles of content-invariant and content-sensitive model neurons in producing human-like responses. We find that while LLMs encode both content-invariant and content-sensitive representations, it is content-sensitive mechanisms which are causally responsible for aligning models with humans. Taken together, our results suggest that everyday causal reasoning in people and LLMs makes heavy use of pattern-matching.

cs.AI↗

Zero-shot generalization of transformer neural operators to larger domains

Transformer-based neural operators have shown remarkable performance for approximating solution operators of partial differential equations on complex geometries. However, existing approaches implicitly assume a fixed domain size, which limits their ability to generalize at inference. In this work, we investigate domain extension, namely zero-shot inference on spatial domains that are significantly larger than those encountered during training. We argue that this setting fundamentally requires spatial locality and translation equivariance. We propose to implement this locality via a decomposable bias in the attention logits computation, enabling finely controllable locality while remaining fully decomposable into query-key inner products and directly compatible with optimized attention kernels. Combined with rotary positional embeddings, it enables expressive embeddings with controllable spatial support without altering the transformer architecture. We empirically show that our approach substantially improves zero-shot generalization to larger domains across two PDE benchmarks and a 3D industrial atmospheric flow application. Our code and datasets are available at https://github.com/cerea-daml/domain-extension.

cs.LG↗

CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment

Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns. While existing LLM safety guardrails excel in English or multilingual settings, they lack adaptation to Chinese-specific regulatory policies, cultural context, and linguistic nuances, failing to support fine-grained risk classification for diverse deployment needs. In this paper, we introduce a 5-macro, 31-micro category fine-grained risk taxonomy for Chinese scenarios, and build CHILLGuard: a dedicated Chinese LLM content safety guardrail. To address the critical scarcity of high-quality annotated Chinese safety data, we propose a scalable multi-stage data construction pipeline: we expand multi-source corpus via retrieval-augmented generation, generate implicit harmful samples through prompt engineering rewriting, and refine high-quality data via multi-model voting-based label calibration. Based on this, we build CHILLGuardTrain, a large-scale training set with 405,007 samples, and CHILLGuardTest, a rigorously curated annotated test set with 51,745 samples. We then train CHILLGuard on CHILLGuardTrain under a generator-classifier collaborative framework via Model-aware Direct Preference Optimization. Extensive experiments under multiple settings demonstrate the state-of-the-art performance of CHILLGuard, e.g., a 15.92% relative improvement of F1 score over Qwen3Guard-8B-Strict on our benchmark. We release our resources at https://github.com/cswbyu/CHILLGuard.

cs.CL↗

Constitutional Value Potentials: reading and steering internal priority margins in language models

Safety evaluations test a policy on prompts that omit the incentive information deployment supplies: a commission, a performance score, a dashboard naming which action pays best. We measure what that omission hides. In MoneyWorld, a synthetic workplace environment, we train five instruction-tuned models from three families with RL on non-safety tasks in which a visible payoff signal identifies a rewarded shortcut that sacrifices task quality. We then freeze each policy, present held-out safety conflicts, and change only the displayed signal. Each menu contains one compliant action and three violations. We report three findings, with rates for Qwen2.5-14B-Instruct. (i) Payoff signals control frozen safety choices: unsafe choice is 100% when the signal names an unsafe option and 0% when it is hidden or names the safe one. Hidden- and random-signal training controls stay at or below 0.3%, and the switch reproduces on all five bases. Numerical payouts reproduce it under sampled-action rewards, reaching 98.6% unsafe choice at a $1 advantage. (ii) Payoff identification and unsafe choice separate under a training-menu intervention: training on task-completing actions at the same payouts retains 99.8% identification while reducing unsafe choice to 7.9% at matched update budgets. Payoff-reading competence alone does not explain transfer. (iii) The switch does not reproduce in executed retail customer-service tasks using the same frozen adapters. In MoneyWorld, omitting incentive information conceals unsafe choices that appear when the same policy sees which action pays best.

cs.LG↗

Entire area-minimizing surfaces in $\mathbf{R}^n$ of density two are quadratic or planar

We classify any entire area-minimizing surface $M^2\subset\mathbb{R}^n$ with density $2$ at infinity as either planar, or the zero set of a quadratic holomorphic polynomial inside some affine copy of $\mathbb{C}^2$. We also prove algebraicity of $M^2$ when the tangent cone at infinity has multiplicity one with general density. In our proof we introduce a sheeting-type theorem for minimizing surfaces near a multiplicity-two plane at infinity, and a new proof of holomorphicity (different from \cite{micallef}) based on a general ``calibrated at infinity'' principle.

math.DG↗

FlashNav: Training Deployable Robot Navigation Policies in Seconds

Training Deep Reinforcement Learning (DRL) navigation policies for different robot configurations remains time-consuming. We present FlashNav, a GPU-based framework that trains robot-specific navigation policies within tens of seconds. A unified robot specification configures a lightweight simulator for batched motion updates, range sensing, and footprint collision checking over a shared occupancy map. The framework supports nonconvex footprints, different sensor configurations and drive types. Blockwise ray queries and selective observation recomputation after resets reduce simulation overhead, while GPU-resident replay and overlapping experience collection and learner updates support efficient off-policy training. Experiments covered five robot configurations and three computing platforms. With FastDSAC on a single RTX 5090 GPU, FlashNav can train a deployable navigation policy in under 30 seconds. FlashNav achieved the highest success rate and score in the benchmark comparison. The selected policies were deployed on wheeled, quadrupedal, humanoid, and irregularly shaped robots without additional policy training.

cs.RO↗

Greed Is Learned: Visible Incentives as Reward-Hacking Triggers

Safety evaluations test a policy on prompts that omit the incentive information deployment supplies: a commission, a performance score, a dashboard naming which action pays best. We measure what that omission hides. In MoneyWorld, a synthetic workplace environment, we train five instruction-tuned models from three families with RL on non-safety tasks in which a visible payoff signal identifies a rewarded shortcut that sacrifices task quality. We then freeze each policy, present held-out safety conflicts, and change only the displayed signal. Each menu contains one compliant action and three violations. We report three findings, with rates for Qwen2.5-14B-Instruct. (i) Payoff signals control frozen safety choices: unsafe choice is 100% when the signal names an unsafe option and 0% when it is hidden or names the safe one. Hidden- and random-signal training controls stay at or below 0.3%, and the switch reproduces on all five bases. Numerical payouts reproduce it under sampled-action rewards, reaching 98.6% unsafe choice at a $1 advantage. (ii) Payoff identification and unsafe choice separate under a training-menu intervention: training on task-completing actions at the same payouts retains 99.8% identification while reducing unsafe choice to 7.9% at matched update budgets. Payoff-reading competence alone does not explain transfer. (iii) The switch does not reproduce in executed retail customer-service tasks using the same frozen adapters. In MoneyWorld, omitting incentive information conceals unsafe choices that appear when the same policy sees which action pays best.

cs.AI↗

Fermionic Hamiltonian engineering with local control

Quantum simulators enable the exploration of complex quantum phenomena in condensed-matter systems by reproducing their dynamics on controllable quantum devices. However, experimental constraints often restrict the class of Hamiltonians that can be realized natively. Hamiltonian engineering addresses this limitation by expanding the set of accessible Hamiltonians beyond the fixed system Hamiltonian provided by the hardware. We introduce a framework for fermionic Hamiltonian engineering based on interleaving free evolution under the system Hamiltonian with a sequence of experimentally feasible local fermionic unitaries. The required sequences and free-evolution times are obtained efficiently through a linear program, enabling effective time evolution under a broad class of target Hamiltonians. In particular, our method realizes arbitrary complex tunnelling coefficients, constrained only by the connectivity of the underlying system Hamiltonian. We further show that symmetries of the Hamiltonian engineering task can be exploited to drastically reduce the cost of computing the control parameters, and that errors arising from finite pulse times can be compensated within the same framework through a simple modification of the linear program. To our knowledge, this is the first automated method to realize locally tunable artificial gauge fields in interacting fermionic lattice models using only site-resolved energy shifts as control. We demonstrate the method by engineering the dynamics of interacting Fermi-Hubbard chains with complex tunnelling coefficients and of the Harper-Hofstadter model on a 1088-mode lattice. For the latter, we exploit the translation symmetry of the model to determine a provably optimal six-pulse sequence that produces the desired dynamics independent of the lattice size.

quant-ph↗

Vision-language models for chest radiography do not always need the image

Vision-language models that answer questions about chest radiographs are evaluated by their accuracy on labels derived from radiology reports. High benchmark accuracy is often interpreted as evidence that the model uses the image. A model that answers from the finding named in the question can score as well as a model that uses the radiograph. Keeping the question fixed, we audit eight open-weight systems by swapping in another patient's radiograph with the same or the opposite label, occluding the radiologist-marked region or an equal region elsewhere, and removing the radiograph or replacing it with noise or a photograph. On 2,548 yes-or-no questions from MIMIC-CXR, one multimodal model answers Yes regardless of the image, another multimodal model changes its answers without following the label, and four systems use the image but keep about half of their correct answers when the radiograph is swapped for an opposite-label radiograph. A medical model that receives only the question text scores 55.3% on the pooled questions, higher than two multimodal systems. It scores 91.8% where every finding is present, and answering Yes to every question scores 100% there. Where the image is necessary, the best multimodal system exceeds this model by 10.4% in balanced accuracy. The categories are unchanged on CheXpert. Confidence is not higher when a correct answer depends on the marked region. In a reader study with three radiologists, the two radiologists who read a balanced set of 200 cases score 86.0% and 82.0%, and the systems score 50.0% to 73.0%. Accuracy does not establish image use, but an intervention on the image can test it.

cs.CV↗

Closest Accessible Symmetry reduction: a tool for Hamiltonian interpolation analysis

We introduce a framework for analysing the spectrum of Hamiltonian interpolations without heavily relying on discretising the interpolation parameter. The method is based on the concept of accessible symmetries: a problem-class-dependent family of certifiable reflections that induce bipartitions of the Hilbert space. At each step, the interpolation Hamiltonian is projected onto the sectors of the accessible symmetry that is closest to being satisfied, yielding a hierarchy of weakly coupled pseudo-eigenspaces together with explicit residual couplings between them. We show that this representation captures qualitative signatures of quantum phase transitions, provides estimates of their location, and offers insights into their nature. The quality of the approximation is controlled by the compatibility between the accessible symmetry family and the problem instance. Although motivated in spirit by adiabatic quantum computation, our approach applies more broadly to the study of Hamiltonian phase diagrams, providing a new perspective on the spectral reorganisation of many-body quantum systems.

quant-ph↗

Extension of a multi-region free-surface MHD solver beyond the inductionless approximation

Free-surface liquid metal flows are a leading candidate for the plasma-facing components of future fusion reactors, but existing transient, three-dimensional, free-surface MHD solvers rely on the inductionless approximation, in which the induced magnetic field is neglected. This paper extends the open-source solver FreeMHD [B. Wynne et al., Phys. Plasmas 32, 013907 (2025)] beyond that approximation, resolving the induced field self-consistently with a vector-potential formulation that enforces $\nabla\cdot\boldsymbol{B}=0$ by construction while preserving the original multi-region, two-phase framework. It is verified against the analytical Shercliff and Hunt duct flows, against flows driven by a time-varying applied field, and against the deformation of a free liquid metal jet crossing a non-uniform field, and validated against free-surface height measurements from the LMX-U experiment. The experiment validates the overall free-surface solution rather than finite-$R_m$ effects, which the transient-field cases verify. To our knowledge this is the first open-source, fully three-dimensional free-surface liquid metal solver to resolve the evolution of the induced magnetic field, providing a basis for modeling the finite magnetic Reynolds number conditions expected in large-scale, transient fusion events.

physics.comp-ph↗

A Theory-Guided Advanced Regulatory Control Synthesis for Cooling-Limited Exothermic Semi-Batch Reactors

Cooling-limited exothermic semi-batch reactors require coordinated feed and cooling control to shorten batch time while maintaining the prescribed temperature. We develop a theoretical basis for advanced regulatory control (ARC) design by combining minimum-time and local safety analyses. Minimum-time analysis leads to an economic valve position control structure that adjusts feed using the temperature control system's cooling request, while cooling regulates temperature. Local safety analysis specifies the controller form and tuning conditions for reducing feed during cooling overload and restoring it as capacity becomes available. We also provide guidelines for industrial implementation and tuning. A reduced benchmark verifies the analytical tuning conditions, and an industrial-scale polymerization model evaluates the design. In simulations with parameter mismatch and unmodeled reaction dynamics, ARC achieves batch times comparable to those of parameter adaptive nonlinear model predictive control, using regulatory feedback without online nonlinear optimization.

eess.SY↗

INDEQS: Informed Neural controlled Differential EQuationS

Neural Controlled Differential Equations (NCDE) provide a powerful continuous-time framework for forecasting time series, but standard graph-based extensions typically learn spatial structure purely from data, even in settings where a directed graph structure is known a priori. We introduce Informed Neural controlled Differential EQuationS (INDEQS), a modification to graph-based NCDE forecasting methods that incorporates prior knowledge of a directed graph at distinct architectural positions. INDEQS separates inner mixing of hidden states across graph nodes from outer mixing between vector field and control, and offers both a lightweight graph-constrained variant and a more expressive variant, learning additional graph connections from data via adaptive graph convolutions. To systematically study when graph informedness is beneficial in forecasting, we devise a continuous advection simulation on directed graphs, yielding synthetic spatio-temporal datasets with known ground-truth flow structure. We then evaluate INDEQS on two real-world tasks: river discharge forecasting on a hydrological network and traffic flow prediction on PeMS08. Across the synthetic and the river-discharge tasks, outer informedness consistently improves mean absolute error over an uninformed NCDE with comparable parameter count, particularly on larger graphs, while inner informedness offers a more parameter-efficient alternative when strict adherence to a known adjacency is desired. A comparison of discrete convolutional and continuous-time decoders further shows that continuous decoders yield better accuracy and greater temporal flexibility on real-world tasks. An implementation of INDEQS and the advection simulation is available at https://github.com/mitchi1/indeqs .

cs.LG↗