Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 325 records · Page 18Linked to original sources

Probing Persona-Dependent Preferences in Language Models

Large language models (LLMs) can be said to have preferences: they reliably pick certain tasks and outputs over others, and preferences shaped by post-training and prompting appear to influence much of their behaviour. But models can also adopt different personas which have radically different preferences. How is this implemented internally? Does each persona use its own preference representations, or are some representations shared? We train linear probes on residual-stream activations of Gemma-3-27B and Qwen-3.5-122B to predict revealed pairwise task choices, and identify a genuine preference vector: it tracks the model's preferences as they shift across a range of prompts and situations, and on Gemma-3-27B steering along it causally controls pairwise choice. Some preference information transfers across the prompted personas we test: a probe trained on the helpful assistant predicts and steers the choices of qualitatively different personas, including an evil persona whose preferences anti-correlate with the Assistant's.

cs.CL↗

Learning Local Constraints for Reinforcement-Learned Content Generators

Constraint-based game content generators that learn local constraints from existing content, such as Wave Function Collapse (WFC), can generate visually satisfying game levels but face challenges in optimizing global properties, such as playability. On the other hand, reinforcement-learning-trained generators can optimize global properties---because such properties can easily be included in reward functions---but the results can be visually dissatisfying. In this paper, we explore ways to combine these methods. Specifically, we constrain the action space of a PCGRL generator with constraints learned by WFC, effectively allowing the PCGRL generator to achieve global properties while being forced to adhere to local constraints. To better analyze how this hybrid content generation method operates, we vary the number and type of inputs, and we test whether to randomly collapse the starting state and exclude rare patterns. While the method is sensitive to hyperparameter tuning, the best of our trained generators produce visually satisfying and playable puzzle-platform game levels---such as Lode Runner levels---with desired global properties.

cs.AI↗

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models

In this paper, we propose GTA-VLA(Guide, Think, Act), an interactive Vision-Language-Action (VLA) framework that enables spatially steerable embodied reasoning by allowing users to guide robot policies with explicit visual cues. Existing VLA models learn a direct "Sense-to-Act" mapping from multimodal observations to robot actions. While effective within the training distribution, such tightly coupled policies are brittle under out-of-domain (OOD) shifts and difficult to correct when failures occur. Although recent embodied Chain-of-Thought (CoT) approaches expose intermediate reasoning, they still lack a mechanism for incorporating human spatial guidance, limiting their ability to resolve visual ambiguities or recover from mistakes. To address this gap, our framework allows users to optionally guide the policy with spatial priors, such as affordance points, boxes, and traces, which the subsequent reasoning process can directly condition on. Based on these inputs, the model generates a unified spatial-visual Chain-of-Thought that integrates external guidance with internal task planning, aligning human visual intent with autonomous decision-making. For practical deployment, we further couple the reasoning module with a lightweight reactive action head for efficient action execution. Extensive experiments demonstrate the effectiveness of our approach. On the in-domain SimplerEnv WidowX benchmark, our framework achieves a state-of-the-art 81.2% success rate. Under OOD visual shifts and spatial ambiguities, a single visual interaction substantially improves task success over existing methods, highlighting the value of interactive reasoning for failure recovery in embodied control. More details of the project can be found here: https://github.com/FutianLabs/GTA-VLA.

cs.RO↗

Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs

When an omnimodal large language model accepts a question whose textual premise contradicts what it actually sees or hears, does the failure lie in perception or in action? Recent omnimodal models are positioned as perception-grounded agents that jointly process video, audio, and text, yet a basic form of grounding remains untested: catching a textual claim that conflicts with the model's own sensory input. We introduce IMAVB, a curated 500-clip benchmark of long-form movies with a 2x2 design crossing target modality (vision, audio) and premise condition (standard, misleading), which lets us measure conflict detection separately from ordinary multimodal comprehension. Across eight open-source omnimodal LLMs and Gemini 3.1 Pro, we document a Representation-Action Gap: hidden states reliably encode premise-perception mismatches even when the same models almost never reject the false claim in their outputs. Behaviorally, models fall into two failure modes: under-rejection, in which they answer misleading questions as if the false premise were true; and over-rejection, in which they reject more often but also reject standard questions, sacrificing ordinary comprehension accuracy. The gap is modality-asymmetric (audio grounding underperforms vision) and prompt-resistant across seven variants. As an initial diagnostic intervention, a probe-guided logit adjustment (PGLA) re-injects the encoded mismatch signal into decoding and consistently improves rejection behavior. Together, these results suggest the bottleneck for omnimodal grounding lies in translation, not perception.

cs.AI↗

Generalized Model Fractional Quantum Hall States on Lattices

Model wavefunctions are essential for studying fractional quantum Hall (FQH) phases. They reveal the intricate connections between model FQH states, conformal field theory, Jack polynomials, and pseudopotential Hamiltonians. However, most model wavefunctions studied so far are in the continuum. In this work, we construct model wavefunctions and the corresponding exact parent Hamiltonians for the entire $Z_k$ Read--Rezayi series on a lattice. Our model wavefunctions obey a generalized clustering rule implemented on the lattice, which we find is encoded in the Macdonald polynomials providing a unified algebraic description of our lattice-specific model states. This ensures that the model states are exact zero-energy ground states of local density interactions and exhibit an infinitely large gap in the particle entanglement spectrum. Numerical studies of the Laughlin and Moore--Read series find the expected zero-mode and one-flux quasihole counting and positive energy gaps at the system sizes studied, and the available finite-size scaling suggests nonzero thermodynamic gaps. Tuning the model interaction range can drive a phase transition out of the FQH phase. Our work not only advances the theoretical understanding of the underlying mathematical structure of lattice model FQH states but also is potentially useful for designing exotic many-body phases in engineered quantum platforms.

cond-mat.str-el↗

Lyman-alpha Radiation Pressure in Dense Star Clusters: Implications for Star Formation and Winds at Cosmic Dawn

Observations with the JWST in lensed fields have revealed that galaxies at cosmic dawn may concentrate their star formation in highly dense, compact, star clusters. The high columns and low metallicities encountered in their birth environments suggest that Lyman-alpha (Ly$α$) radiation pressure may be crucial to their formation and evolution. In this study, we address this question by post-processing snapshots from radiation hydrodynamic simulations of dense star cluster-forming clouds ($Σ_*\gtrsim10^3{M_\odot{pc}^{-2}}$) with a range of dust abundances ($Z_d=0-0.1Z_{d,\odot}$) using the COLT Monte Carlo code. We infer that Ly$α$ is likely to have mild (~10%) effects on the gas-to-star conversion efficiencies ($ε_*\gtrsim60$%) for $Z_d\gtrsim0.01Z_{d,\odot}$, and even in dust-free environments, $ε_*\gtrsim25$% - much higher than the <10% values typical of star-forming regions in the local Universe. This is because the densest filaments dominating stellar mass assembly ($n\gtrsim10^4{cm}^{-3}$) remain sub-Eddington ($f_{Edd}<1$). On the other hand, the bulk of the gas volume ($n\lesssim10^3{cm}^{-3}$) has $f_{Edd}>1$, with noticeable fractions having $f_{Edd}\gtrsim10$, implying that Ly$α$ can launch dynamically significant winds from these systems rapidly ($\lesssim$4Myr), with possible implications for ionizing photon escape and galactic outflows. The Ly$α$ force multiplier $M_F$ is highly sensitive to $Z_d$, with $M_F\lesssim3$ ($\lesssim 500$) for $0.1Z_{d,\odot}$ (dust-free) environments respectively. Nevertheless, Ly$α$ dominates over UV and IR radiation pressure at all values of $Z_d\lesssim0.1Z_{d,\odot}$, by factors of ~3-500. Our results suggest that Ly$α$ radiation pressure reinforces the emerging picture of locally efficient, bursty star formation accompanied by rapid outflows in galaxies at cosmic dawn.

astro-ph.GA↗

Computing Lower Bounds on the Nonnegative Rank via Non-Convex Optimization Solvers

The nonnegative rank of a nonnegative matrix $X$ is the smallest number of nonnegative rank-one factors that sum to $X$. Since computing the nonnegative rank is NP-hard, it is common to circumvent this issue by computing lower and upper bounds. In this paper, we propose non-convex formulations and practical implementations for four important lower bounds for the nonnegative rank, namely the fooling set bound (FSB), the rectangle covering bound (RCB), the hyperplane separation bound (HSB), and the self-scaled bound (SSB). In particular, our algorithm for computing the SSB is the first available in the literature, to the best of our knowledge. It allows us to improve the best known lower bound on the nonnegative rank for some matrices. In some cases, they coincide with the best known upper bound, thereby establishing their exact nonnegative rank for the first time. Moreover, on canonical benchmarks, we show that our non-convex approaches provide a meaningful and often competitive alternative to standard methods. The paper also provides a consolidated reference for the current state of several classical lower bounds on a large number of benchmark matrices.

math.OC↗

GPart: End-to-End Isometric Fine-Tuning via Global Parameter Partitioning

Low-rank adaptation (LoRA) has become a dominant paradigm for parameter-efficient fine-tuning (PEFT) of large-scale deep learning models. However, its bilinear parameterization induces a parameter-dependent geometry: the mapping from trainable parameters to weight updates is not generally distance-preserving. Related methods that project a low-dimensional vector into LoRA's parameter space, such as Uni-LoRA, improve parameter efficiency, but the subsequent bilinear map breaks end-to-end isometry. We propose GPart (Global Partition fine-tuning), a highly parameter-efficient fine-tuning method that maps a $d$-dimensional trainable vector directly into the full weight space through a sparse, isometric partition matrix. GPart retains a fixed global parameter-sharing prior while removing the additional low-rank reconstruction used by LoRA-based methods. This yields a simple parameterization with a single main hyperparameter ($d$), exact end-to-end isometry, and a minimal checkpoint representation consisting of the trainable vector and a random seed. GPart builds on the premise of effective fine-tuning within random low-dimensional subspaces of the full weight space without requiring a low-rank matrix factorization. Across natural language understanding, computer vision, and mathematical reasoning benchmarks, GPart matches or improves over existing PEFT methods at ultra-low parameter budgets. Beyond offering mathematical tractability and memory efficiency, the direct linear parameterization of GPart streamlines model selection and paves the way for compact adapter composition. Overall, GPart provides an elegant and competitive alternative for fine-tuning under small parameter budgets, with a fixed and predictable geometry between trainable coordinates and weight-space updates.

cs.LG↗

Learned Suppression for 3D Keypoint Detection with a Graph-Transformer Backbone

Detecting 3D keypoints is a long-standing challenge in computer vision. Most detectors end with a heuristic post-processing step that is not learned. We propose a 3D keypoint detector that improves on this step with a learned suppression module, paired with a Point Transformer backbone that we extend with a directional graph neural network. The module is a graph network over candidates that learns which to keep, which to suppress, and how to relocate the remaining ones. Paired with three backbones, it improves over DBSCAN and greedy non-maximum suppression, and because it operates on candidate features rather than raw geometry, the same formulation applies to both structural and semantic keypoints. Our model surpasses the per-category trained KeypointDETR on 12 of 16 KeypointNet categories, attains the best Corner F1 on the Building3D Entry-Level benchmark, and remains competitive with BWFormer on the larger Tallinn split. GitHub implementation: https://github.com/cansdev/learned-suppression-3d.

cs.CV↗

Universal Approximation of Nonlinear Operators and Their Derivatives

We show that Universal Approximation (UA) of nonlinear operators and their derivatives via Operator Learning (OL) architectures fails in ${C^k_F}$ (Fréchet) compact-open topologies and in Fréchet--Sobolev norms (i.e. under operator norms). We solve this obstruction by restoring UA in natural weaker topologies: $C^k_B$ (Bastiani) compact-open topologies and (novel) weighted Bastiani--Sobolev spaces for general finite input measures. In full Banach-space generality, these are the first complete generalizations of the corresponding influential classical results in [Hornik, 1991] to infinite-dimensional spaces and OL. Based on our UATs, we formulate Bastiani--Sobolev training in DIOL. These results launch Derivative-Informed Operator Learning (DIOL) (i.e. learning nonlinear operators and their derivatives) on general Banach spaces. We parameterize nonlinear operators via Encoder-Decoder Architectures, classical OL architectures available in general Banach spaces; these include DeepONets, Deep-H-ONets, and PCA-Nets, which our UATs cover. A key mathematical result is that our new weighted Bastiani--Sobolev spaces generalize classical Gaussian (Malliavin) Sobolev spaces on Banach spaces. Open frontiers where DIOL and our UATs find applications are: high-order accuracy in OL; fast constrained optimization in Banach spaces (e.g. optimal control of PDEs, inverse problems) via Learn-Then-Optimize; numerical methods for infinite-dimensional PDEs (e.g. HJB PDEs on Banach spaces from infinite-dimensional optimal control via Optimize-Then-Learn, such as optimal control of PDEs, SPDEs, path-dependent systems, partially observed systems, mean-field control).

cs.LG↗

Reshape and Recur: Improving SSMs with Input Reshaping and Depth Recurrence

State Space Models (SSMs) are increasingly deployed in the Edge because they offer, at comparable performance, a smaller memory/training/inference footprint, compared to Large Language Models (LLMs). These three advantages are a direct consequence of the time recurrence inherent in the SSMs architecture. Here, we further improve this recurrent architecture by positively answering two previously underexplored, orthogonal questions: (1) Can we reduce SSMs memory-footprint without any performance penalty, by also employing depth recurrence? (2) Can we increase SSMs performance by using a fixed and consistent time-granularity across all tasks? The first question is somewhat unexpected, given that SSMs are already recurrent. However, the orthogonal depth recurrence further decreases SSMs memory footprint. We show that a looped SSM with $k$ parameters adaptively iterated $M$ times, achieves a performance comparable to a standard SSM with $k \cdot L$ independent parameters, where $M \leq L$. The second question is also unexpected given the time-recurrent nature of the SSMs architecture. However, it makes perfect sense for the time-parallel training of SSMs on the entire input sequence. We show that concatenating time steps for lower-dimensional sequence elements, or flattening and re-chunking the joint feature-time dimension for high-dimensional ones, can improve the baseline by enhancing the way information is presented to the model. Our results for both extensions lead to consistent benefits across four representative SSM architectures: LRU, S5, LinOSS, LrcSSM.

cs.LG↗

An agitated oscillator chain

Reduced dynamical descriptions for spatially extended systems in contact with a nonequilibrium medium give essential information about extensions of fluctuation-dissipation relations and, at the same time, are relevant for many natural systems. Here we study a chain of harmonic oscillators coupled to fast run-and-tumble particles, discretizing earlier work on an elastic string. We derive the effective Langevin dynamics, including the explicit discrete counterparts for the nonequilibrium-induced streaming term, friction coefficient, and noise amplitude. At sufficiently high bath persistence, the linear friction becomes negative, destabilizing the chain dynamics. However, numerical simulations confirm that the anti-damping is eventually arrested and results in a nonlinear stationary condition, whose properties are the central topic of this work. We characterize the long-time phase-space trajectories and their Rayleigh-like structure, identify multistage relaxation of the mean-squared displacement and mean-squared velocity, and establish a non-monotonic dependence of stationary fluctuations on bath persistence. We further analyze the stationary distributions of the displacement and velocity, as well as the spatial correlations, which develop damped oscillations in the strongly persistent regime. The transfer of activity thus transforms a passive harmonic chain into a self-sustained fluctuating medium with many-body Rayleigh-like dynamics, pulsating displacements, and persistent velocities.

cond-mat.stat-mech↗

On ultraproduct approximations and property (T) factors

We introduce a framework allowing for key aspects of deformation/rigidity theory to be used in the study of continuous model theory of II$_1$ factors. Using this framework, we solve several well-known open problems in the area. For example, we show that the group von Neumann algebras $L(\operatorname{SL}_3(\mathbb Z))$ and $L \mathbb F_2$ are not elementarily equivalent, and we show that the group von Neumann algebra $L\mathbb F_2$ is not pseudomatricial. We also show a Bass-Serre type strong rigidity result in the setting of ultraproducts to provide an infinite family of pairwise non-elementarily equivalent full factors, each of which embeds into an ultraproduct of the hyperfinite II$_1$ factor. Building on previous work of Boutonnet, Chifan and Ioana, we also provide a continuum of pairwise non-elementarily equivalent full factors, which can be taken to be group von Neumann algebras.

math.OA↗

Graph Hierarchical Recurrence for Long-Range Generalization

Graph Neural Networks and Graph Transformers have become central to graph learning, combining expressive representation learning with sample-efficient inductive biases. Yet they remain fundamentally limited when predictions depend on correlations between distant graph regions. We address this limitation with Graph Hierarchical Recurrence (GHR), a novel framework that jointly operates on the input graph and a pooled hierarchical abstraction. We also show that existing models degrade more sharply under out-of-range generalization, where test instances require interactions across distances exceeding those observed during training. Despite its minimal design, GHR consistently strengthens every tested message-passing backbone, yielding robust performance on long-range dependencies and particularly pronounced gains in out-of-range regimes. Across a broad suite of long-range benchmarks, GHR achieves state-of-the-art or competitive results on multiple tasks, establishing hierarchical recurrence as an effective mechanism for extending graph models beyond their observed interaction range.

cs.LG↗

DexHoldem: An Agentic Robotics Benchmark for Dexterous Manipulation in Texas Hold'em

Evaluating embodied systems with real dexterous hardware requires more than isolated motor-skill tests: an agent must perceive a changing scene (e.g. a tabletop), choose a context-appropriate action, execute it with a dexterous hand, and leave the scene usable for later decisions. We introduce DexHoldem, a comprehensive real-world benchmark evaluating Texas Hold'em related dexterous manipulations with a ShadowHand. DexHoldem provides 1,470 teleoperated demonstrations across 14 Texas Hold'em manipulation primitives, a standardized physical policy benchmark, and an agentic perception benchmark that tests whether agents can recover the structured game state needed for embodied decision making. On primitive execution, $π_{0.5}$ obtains the highest task completion rate ($61.2\%$), while $π_{0.5}$ and $π_0$ tie on scene-preserving success rate ($47.5\%$). On agentic perception, Opus 5.5 narrowly leads on both strict problem-level accuracy ($49.1\%$) and average field-wise accuracy ($80.6\%$); the gap between the two exposes the distance between isolated visual sub-capabilities and complete routing-relevant state recovery. Finally, we instantiate the full embodied-agent loop with one agent--policy pairing over 33 closed-loop hand-level rollouts, in which only $12.1\%$ of hands complete; retries restore the failed primitive in 12 of 34 dispatches and resolve prolonged execution stalls in three of the four completed hands, which would otherwise have required manual termination. Only one hand completes with neither a retry nor a human-help request. DexHoldem therefore evaluates dexterous tabletop execution, agentic perception, and embodied decision routing in a shared physical setting. Project website: https://dexholdem.github.io/Dexholdem/

cs.RO↗

4D and 5D Layer Codes through Color Routing

We introduce and explicit Calderbank-Shor-Steane (CSS) code construction that generalizes the Layer codes to $D=4,5$ dimensions. Much like its predecessor, the present construction is based on embedding quantum low-density parity check (qLDPC) codes; from an $[[n,k,d]]$ code with energy barrier $Δ$, we obtain a $D=4,5$ dimensional Layer code with parameters $[[Θ(n^{D/(D-2)}), k, Θ(dn^{1/(D-2)})]]$ and energy barrier $Ω(Δ)$. Using good qLDPC codes as input, our construction saturates the $D=4,5$ dimensional BPT bounds exactly. The higher dimensional Layer Codes are modular, and thus well suited to architectures composed of modular network patches, despite our physical limitation to three dimensions. We overcome the hurdles encountered by previous generalization attempts through the use of \textit{color routing}, allowing us to resolve the structure of the check layers and line defects.

quant-ph↗

The Markov Marginal Problem for Density Operators

We study when local reduced density operators, viewed as quantum marginals, can be assembled into a global quantum state with a prescribed Markov structure. The starting point is a canonical logarithmic construction $T(\mathcal R)$, the noncommutative analogue of the junction-tree formula for decomposable graphical models. Unlike in the classical case, this formal construction may fail: noncommutativity can prevent it from being a normalized state with the prescribed marginals. We prove that this obstruction is captured exactly by a trace condition. For two overlapping marginals, and for clique marginals on a chordal graph, the condition ${\rm Tr}(T(\mathcal R))=1$ is equivalent to the existence of a quantum Markov completion. When it exists, the completion is unique, equal to $T(\mathcal R)$, and selected by the maximum entropy principle. In the two-clique case, we also give an equivalent conditional reconstruction characterization: the two natural one-sided sandwich reconstructions agree if and only if the trace condition holds. We introduce the global quantum information $g{\rm I}(\mathcal{G})_ρ$ associated with a chordal graph $\mathcal{G}$ and show that it is a relative-entropy discrepancy from $ρ$ to the logarithmic candidate, with a trace correction when the candidate is not normalized. We also prove an intersection property for strictly positive quantum conditional independence. Three-qubit Pauli examples illustrate how the quantum obstructions are real: local consistency, feasibility, Markov feasibility, and maximum entropy can all separate.

quant-ph↗

An Energy Integration Free Kubo-Bastin Formula Decomposition

Kubo formulae play a central role in modern spintronics and condensed matter physics, serving as the foundational ground for studying transport responses in the linear regime. In this work, we propose a reformulation of the widely used Kubo-Bastin decompositions that eliminates the need for numerical energy integration. By performing these integrations analytically for generic periodic systems, our approach significantly reduces computational cost and simplifies the calculation of transport coefficients, while maintaining a clear distinction between the relevant Fermi-surface and Fermi-sea transport components. Our formulation suggests that the Fermi-surface term can be attributed to a pure intraband contribution. In contrast to widely used energy integration free expressions that are strictly limited to zero temperature, the present formulation remains valid at arbitrary temperatures.

cond-mat.mes-hall↗