Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

IR-SIM: A Lightweight Declarative Simulator for Navigation Learning and Benchmarking

Developing navigation policies requires simulation scenarios that support repeatable training and evaluation. Despite the availability of numerous simulators, constructing diverse navigation scenarios often involves writing simulator-specific code, using complex graphical interfaces, performing substantial manual configuration, or relying on resource-intensive computing platforms, making scenarios difficult to reproduce and reuse across experiments. To this end, we develop the Intelligent Robot Simulator (IR-SIM), a lightweight declarative simulator implemented as a Python library to support navigation learning and benchmarking through rapid construction of reusable and diverse scenarios and efficient execution on accessible platforms. IR-SIM represents scenarios as human-readable YAML configurations, in which necessary objects, behaviors, sensors, maps, and environment parameters can be flexibly specified and composed. This declarative representation makes IR-SIM friendly to large language model (LLM)-powered agents, enabling them to compose and modify executable navigation scenario files through the provided agent skills, instead of writing a large amount of code that is difficult to reuse. Despite its lightweight implementation, IR-SIM provides the essential components for navigation simulation, while retaining interfaces to high-fidelity simulators for downstream validation. Experiments demonstrate that IR-SIM runs up to several hundred times faster on the evaluated CPU platform, while agent skills reduce mean LLM-based scenario construction time by more than $50\%$ across both evaluated models. The experiments further show ability to support reproducible benchmarking and RL-based navigation policy learning.

cs.RO↗

Generalizing Geometry-Guided Mamba as a Plug-and-Play Context Module for CNN-based Semantic Segmentation

CNN-based semantic segmentation networks usually rely on context heads such as ASPP, PPM, or attention modules to enlarge the receptive field. These heads are effective but may introduce heavy computation, memory cost, or boundary leakage. This paper revisits Directional Geometric Mamba (G-Mamba) from DGM-Net and studies it as a plug-and-play context aggregation module rather than a completely new segmentation architecture. The key idea is to inject geometric guidance into the selective scan process, allowing long-range feature propagation to be modulated by boundary and centripetal-flow cues. We replace the original context heads of six representative CNN segmentation models, including DeepLabV3+, DANet, CCNet, PSPNet, PSANet, and OCRNet, while keeping the ResNet-101 backbone unchanged. On CCNet, we additionally compare serial and parallel combinations of criss-cross attention and the G-Mamba block, with the parallel head performing best. Results on Cityscapes show consistent mIoU gains with only moderate extra GFLOPs at $1024\times1024$ resolution, suggesting that geometry-guided SSM modules can serve as practical alternatives or enhancements to conventional CNN context heads.

cs.CV↗

Algebraic Kolmogorov--Arnold representation theorem for quantum measurement

We establish an operational framework connecting the classical Kolmogorov--Arnold (KA) representation theorem to quantum information theory. By introducing and proving an algebraic, bounded-degree polynomial version of the theorem, we demonstrate that any target physical property of an unentangled multi-qubit product state can be exactly decomposed using a finite, fixed set of local ``inner'' observables and a shallow architecture of univariate polynomials. We further analyse the stability of this Quantum Kolmogorov--Arnold (QKA) representation under adversarial perturbations. In stark contrast to the pathological instabilities and severe reparameterisation sensitivities inherent to the classical Kolmogorov--Arnold representation theorem, our algebraic quantum framework exhibits remarkable resilience. We prove that the representation remains stable against bounded physical perturbations acting on the inner measurement operators, and show via the Heisenberg picture that it is inherently immune to adversarial quantum channel attacks acting on the input states.

quant-ph↗

UPLOTS: A Unified Pretrained Language Model for Constrained Time-series Generation

In time-series generation, existing approaches typically handcraft ortrain a separate model for each dataset, which hinders their scalability and fails to leverage shared temporal structures across domains. To address this fragmentation, we propose UPLOTS, a Unified, Prompt-guided Language model framework fOr constrained Time-Series Generation across diverse domains. Instead of building task-specific models, UPLOTS leverages a single pre-trained transformer backbone guided by learned constraint prompts, enabling on-demand generation with precise pattern control. One key innovation is our dynamic multi-dataset loss re-weighting and prompt-to-pattern mapping, which allows UPLOTS to internalize diverse temporal structures during training and conditionally generate them at inference. We evaluate UPLOTS on four real-world benchmarks and multiple constraint settings, including peak-period, calendar, load-level, and volatility patterns. Additional held-out constraint-combination and downstream forecasting experiments further demonstrate that UPLOTS generalizes beyond the original peak-pattern setting and improves data augmentation under scarce real-data regimes. Our code and baselines are available at github repo: https://github.com/cruiseresearchgroup/UPLOTS.

cs.LG↗

Space-sampled Value Decay: Forgetting Mechanisms for Non-stationary Reinforcement Learning

Reinforcement Learning agents deployed on physical systems must adapt continually, since degradation and shifting environment conditions change the dynamics (they \emph{drift}) over time. In the hardest version of this problem, the agent interacts with a single system that might drift at every timestep, leaving no opportunity to revisit past conditions -- a setting we call Single Environment, One-Shot Non-Stationary Reinforcement Learning (SEOS-NSRL). We argue that this setting calls for selective forgetting rather than re-learning, and introduce Space-sampled Value Decay (SsVD), which pulls value estimates of randomly chosen elements of the state space to a baseline value, so that outdated information in non visited regions is discarded. SsVD does not require resetting or change-point detection and plugs into modern off-policy algorithms; we integrate it into Soft Actor Critic and Deep Q-Networks. Across 6 non-stationary environments, SsVD improves upon its direct base algorithms and attains the best mean rank across all. The SsVD mechanism can also induce optimism which we show on hard-exploration tasks, although we investigate the connection here only briefly.

cs.LG↗

JointEdit3D: Feed-Forward 3D Scene Editing in a Unified Latent Space

Existing 3D scene editing methods typically rely on per-scene optimization over explicit 3D representations or cascaded edit-and-reconstruct pipelines, resulting in high test-time cost, limited 3D awareness, and structural inconsistencies. To couple appearance synthesis with geometry prediction, we adapt a pretrained unified RGB-geometry latent space to feed-forward scene editing. Given a source video and an edited reference image, JointEdit3D performs asymmetric latent inpainting: it observes only the edited RGB reference latent and jointly generates the remaining RGB latents and the entire geometry latent along the source trajectory. JointEdit3D introduces a dedicated SceneAnchor Branch to inject source-scene structure without forcing direct copying, and adopts edit/background-aware losses to balance edited-region fidelity with unedited-content preservation. To address the lack of paired resources for standardized 3D scene editing evaluation, we introduce SceneEdit3D-15K, a dataset with 15K paired editing samples and renderer-provided 3D annotations, together with SceneEdit3D-Bench, a curated 100-sample benchmark. Experiments show that JointEdit3D improves edited-region quality and 3D structural completeness over prior baselines while maintaining competitive background preservation.

cs.CV↗

Optimal classical shadow estimation of unitary channels at Heisenberg limit

Full tomography of an unknown quantum evolution is resource-intensive and often unnecessary when the goal is only to predict selected properties. This motivates the study of classical shadow estimation of unitary channels (CSEU), a task in which one queries an unknown $d$-dimensional unitary $U$ and stores classical data that can later be used to predict expectation values $\mathrm{tr}[O UρU^\dagger]$ for arbitrary input states $ρ$ and observables $O$. Here we propose a parallel and non-adaptive CSEU protocol that uses $\mathcal{O}(\min\{d\varepsilon^{-2},d^{3/2}\varepsilon^{-1}\})$ queries to $U$ when the observables have constant effective size. This achieves Heisenberg scaling with respect to the additive error $\varepsilon$ and is query-optimal, as we prove a matching lower bound $Ω(\min\{d\varepsilon^{-2},d^{3/2}\varepsilon^{-1}\})$ for the query complexity, which remains valid even when stronger access to the unknown unitary is available. Our CSEU protocol is a versatile primitive with applications to several quantum learning tasks. In particular, it yields a near-optimal algorithm for learning an arbitrary Hamiltonian from its real-time evolution in the high-precision regime and recovers the best-known bounds for several tasks in a unified way.

quant-ph↗

Elastic ODYN: Differentiable Optimization for Infeasible Control and Learning in Robotics

Robotic systems routinely encounter conflicting objectives, modeling errors, and degenerate contact conditions that render quadratic programs (QPs) infeasible. Yet most optimization solvers and differentiable QP layers assume feasibility, leading to numerical failures, unstable gradients, or solver breakdown when constraints cannot be simultaneously satisfied. We present Elastic ODYN, a primal-dual non-interior-point QP solver that handles infeasibility through smooth squared-$\ell_2$ elastic relaxations. The formulation remains well posed under ill-conditioning and degeneracy, supports warm starting, and converges to closest-to-feasible solutions, with lightweight refinement recovering physically meaningful dual variables. Building on this framework, we develop Elastic ODYNLayer, a differentiable QP layer with stable gradients under infeasibility, and Elastic OdynSQP, an SQP method that resolves inconsistent subproblems and intrinsically infeasible optimal control tasks through selective constraint elasticity. Across benchmark QPs, singular contact mechanics, differentiable parameter identification, and quadrupedal and humanoid trajectory optimization, Elastic ODYN outperforms state-of-the-art elastic QP solvers in robustness, warm-start performance, and convergence reliability, enabling optimization, simulation, control, and learning beyond standard feasibility assumptions.

cs.RO↗

Bounds for the ratio between the independent domination number and the domination number

In this article we present new and improved results for the ratio between the independent domination number and the domination number in graphs with bounded degree. We present a general formula, that, for a fixed maximum degree, allows to compute an upper bound for this ratio as a function of an upper bound $β|V|$ for the independent domination number. We also apply this formula to several known upper bounds for the independent domination number. Furthermore we present constructions giving lower bounds for the best possible upper bound in various classes of graphs with bounded degree.

math.CO↗

ActiveSAM: Fast and Accurate Open-Vocabulary Semantic Segmentation with Frozen SAM 3

Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabulary semantic segmentation (OVSS) is inefficient: full-resolution decoding is typically run over the entire dataset vocabulary, whereas each image contains only a small active subset of classes. We introduce ActiveSAM, a training-free inference framework that turns SAM 3 into an active-vocabulary segmenter. ActiveSAM first canonicalizes and expands class prompts, then uses evidence-proportional grounding to estimate an image-conditioned active set from a low-resolution presence preview. Only retained prompts receive full-resolution mask prediction, using bucketed prompt multiplexing with the frozen SAM 3 decoder. The preview stage uses only class-presence evidence and skips unnecessary segmentation-head computation. To resolve overlapping concept responses, exclusive concept decoding compares each pixel's joint score vector with class signatures estimated once per vocabulary from unlabeled images. ActiveSAM requires no weight updates, no oracle class-presence labels and no per-dataset hyperparameter tuning. Across eight OVSS benchmarks, ActiveSAM improves the speed-accuracy tradeoff of training-free open-vocabulary semantic segmentation, outperforming the current state-of-the-art SegEarth-OV3 by +2.1 mIoU on average while running much faster, with 7.3-12.2x speedups on large-vocabulary datasets. ActiveSAM also achieves the highest accuracy under image corruptions that simulate real-world distribution shift, making it well-suited for deployment in noisy-input domains such as autonomous driving and embodied AI. Code is available at https://github.com/VILA-Lab/ActiveSAM

cs.CV↗

Non-negative Matrix Factorisation with Topological Regularisation

Non-negative matrix factorisation (NMF) learns additive representations from data, but non-negativity alone does not ensure interpretable basis functions. We introduce Top-NMF, which guides basis learning through general topological preferences without prescribing the detailed form of the components. Observations and basis vectors are treated as non-negative functions on structured domains. Discrete topological conditions are difficult to optimise directly. Our guiding viewpoint is that persistent homology provides a stable, continuous relaxation of ordinary homological invariants: it tracks connected components and loops across thresholds, allowing structural preferences to be expressed through continuous scores suitable for optimisation. We construct such scores for connected image parts, clique-like graph structure, and periodic time-series components, and incorporate them into a common NMF objective. For graph data, we derive an exact maximum-weight-spanning-tree formula and characterise all local minimisers of the score, showing that their positive-weight edges form disjoint cliques. We establish the regularity of the scores, describe their derivatives, and analyse projected first-order optimisation. Numerical studies on synthetic and real data demonstrate how these priors guide basis learning and examine the balance between structural agreement, atom recovery, and reconstruction accuracy.

cs.LG↗

Congruences of shifted Jack Littlewood-Richardson coefficients

The shifted Jack Littlewood-Richardson coefficients generalize the ordinary Jack coefficients and are Laurent polynomials in the Jack parameter $α$. We prove a previously conjectured congruence: coefficients indexed by triples differing by a single box move are congruent modulo the shared $α$-hook at the pivot. We also prove a shifted Macdonald analogue, with a power-of-$t$ twist, and establish that the normalized shifted Macdonald coefficients are Laurent polynomials in $q$ and $t$. The proofs combine coincidences of shifted coordinates with Laurent-preserving shift transforms, and the Macdonald input uses Knop's inversion formula and the integrality of the Bergeron-Garsia-Haiman-Tesler operators. Finally, we realize the Jack and Macdonald congruences as necessary edge conditions on Hilbert schemes of points, in equivariant cohomology and equivariant $K$-theory, respectively. In the Macdonald case a tautological determinant twist accounts for the power-of-$t$ normalization

math.CO↗

The Chirp-Mass Ladder: A New Rung Emerges

We study the binary black hole (BBH) population observed through gravitational waves (GWs) using 256 events from the latest release of GWTC-5.0. The inferred chirp-mass distribution shows prominent peaks at approximately $7.5M_{\odot}$, $14M_{\odot}$, and $27M_{\odot}$, with subsequent peaks spaced by approximately a factor of two. A parsimonious explanation for this structured distribution is a hierarchical merger scenario, in which the first peak arises from mergers of black holes of stellar origin, and higher-mass peaks arise from repeated mergers. Notably, with the addition of new observations, an intermediate peak near $19M_{\odot}$ emerges. This feature was anticipated in earlier work as a consequence of intergenerational mergers involving second- and third-generation (G) black holes, highlighting the predictive expectations of the hierarchical-merger interpretation. Furthermore, two groups of $1G+2G$ mergers recently reported in separate studies can be understood as distinct rungs---$1G+2G$ and $3G+4G$---within this hierarchical chirp-mass ladder, a unification that describes both spin transitions with a single mechanism. Although we observe expected correlations between mass ratios and spins in multiple events across the mass range, the lack of clear signatures across all rungs invites investigation into the role of hierarchical mergers in shaping the BBH population.

astro-ph.HE↗

Guava: Distilling Frontier VLMs into a Compact Agent through a Robotic Manipulation Harness

Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. Harnessing models through embodied tools use offers a promising alternative to end-to-end vision-language-action systems by combining high-level reasoning with external modules for perception, planning, and control. However, it remains unclear what makes an effective harness for embodied manipulation, and to what extent such a harness can unlock embodied capabilities in a wide range of reasoning models. In this work, we present Guava, a harness framework for embodied tool use developed through systematic exploration of the design space of agent workflows, action spaces, and observation spaces. Our study identifies three key ingredients for effective embodied agents: iterative perception-reasoning-action loops, semantic action abstractions, and multimodal observations. To understand whether these design principles are universal even to small models, we develop an end-to-end training pipeline that distills embodied manipulation capabilities into a 4B open-source model using fewer than 2K trajectories collected entirely in simulation. Experimental results in both simulation and real-world environments show performance comparable to frontier proprietary models while exhibiting strong generalization to unseen objects, novel instructions, and long-horizon tasks. Results suggest that a well-designed harness can serve as a scalable, model-agnostic interface for embodied manipulation, enabling strong emergent embodied capabilities in compact open-source models with minimal training data.

cs.RO↗

Do Gaussian Scenes Contain Enough Structure for Intrinsic Segmentation?

Gaussian segmentation is usually posed as transferring object knowledge from 2D foundation models into a 3D representation. This leaves a fundamental question unanswered: how much object structure is already encoded by a trained gaussian scene? We investigate this question with GS-IntSeg, an intrinsic, mask-free, and training-free method that constructs partitions using only: gaussian geometry, opacity, spherical-harmonic radiance, and deformation trajectories. On the dynamic Neu3D and HyperNeRF datasets, GS-IntSeg obtains a mean of 0.677 mIoU across multi-view and monocular scenes without masks, external features, or segmentation training. This intrinsic formulation also enables GS-IntSeg to require approximately 2.5 minutes per HyperNeRF scene on a consumer RTX 5080 to construct its partitions, over 10x faster than SAM-based methods that require mask generation, feature rendering, and other stages. These results suggest that gaussians alone can approach mask-supervised performance in gaussian scenes segmentation. While a gap remains between intrinsic and foundation model based methods, to our knowledge, GS-IntSeg is the first mask-free approach to gaussian segmentation, motivating a re-evaluation of the assumption that gaussian scene segmentation must be based on external masks and pointing toward an alternative faster, more generalizable segmentation approach.

cs.CV↗

Skill-MAS: Evolving Meta-Skill for Automatic Multi-Agent Systems

Large Language Model (LLM)-based automatic Multi-Agent Systems (MAS) generation has become a crucial frontier for tackling complex tasks. However, existing methods face a dilemma between model capability and experience retention. Inference-time MAS leverages frozen frontier LLMs but repeats identical searches without learning from past experience. Conversely, Training-time MAS internalizes experience via gradient updates but is constrained by the low capability ceiling of smaller models, and is hard to scale to large frontier LLMs. To bridge this gap, we propose Skill-MAS, a novel third path that decouples experience retention from parametric updates by conceptualizing the high-level orchestration capability as an evolvable Meta-Skill. Skill-MAS refines this architectural knowledge through a closed optimization loop: (1) Multi-Trajectory Rollout samples a behavioral distribution for each task under the current Meta-Skill; and (2) Selective Reflection adaptively selects priority tasks and applies hierarchical contrastive analysis to distill systemic experience into generalizable, strategy-level principles. Extensive experiments across four complex benchmarks and four distinct LLMs demonstrate that Skill-MAS not only achieves remarkable performance gains but also maintains a favorable cost-performance trade-off. Further analysis reveals that the evolved Meta-Skills are highly robust and exhibit strong transferability across unseen tasks and different LLMs.

cs.MA↗

Monocular 3D Occupancy Perception for Robots on Sidewalks via Hybrid 2D-3D Learning

Sidewalks in the real world are crowded, cluttered, and less structured than roads, making 3D occupancy prediction a key ingredient for the safe navigation of mobile robots such as delivery bots and electric wheelchairs. Existing occupancy learning pipelines are largely designed for on-road autonomous driving and often train on large-scale paired LiDAR-RGB datasets with dense 3D supervision and multiple camera inputs, which are costly to collect and do not adequately capture sidewalk-specific characteristics. We propose WalkOCC, a hybrid Ray-marching monocular 3D occupancy perception framework for robots operating on sidewalks. WalkOCC explicitly couples geometric grounding from LiDAR-RGB paired data with scalable learning from large-scale unpaired monocular images. It bootstraps pseudo occupancy supervision from paired sequences and jointly learns image-level representations on additional 2D-only data. It yields stable optimization and improved generalization without requiring costly 3D occupancy annotations. Extensive experiments demonstrate consistent gains in prediction accuracy, fine-grained segmentation of subtle urban structures such as curbs and gutters, and robustness to environmental and cross-embodiment shifts compared with self-supervised image-based baselines. To facilitate evaluation and benchmarking, we also introduce Sidewalk3D, a large-scale sidewalk perception dataset with LiDAR-camera paired sequences collected across multiple locations and time periods, along with 3D semantic occupancy annotations for evaluation. Code and data will be made available.

cs.RO↗