Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

Tool-Policy Co-Design for Powder Weighing in Laboratory Automation

Autonomous powder weighing is one of many bottlenecks in laboratory automation due to the complex, non-linear dynamics of heterogeneous materials. Robot chemists performing this task utilise standard tools shaped for the dexterity of human hands, whose fixed geometry sets the dynamics that the control policy needs to regulate. This work introduces a tool-policy co-design framework that concurrently optimises the morphology of a dispensing tool and its control policy for use by robots in chemistry laboratories, formulated as a bi-level optimisation that minimises dispensing error over a target distribution of powder flowabilities. The outer loop varies tool-design parameters such as tool depth, width and rim spike topology using Bayesian optimisation and hyperband, while an inner loop optimises a control policy for each candidate morphology. We also introduce a geometric similarity metric that warm-starts policy training from cached policies of structurally similar designs, exploring 28% more configurations under the same compute budget. The proposed framework is evaluated on a robotic powder weighing task across seven materials with distinct physical dynamics in a flowability-informed robot-material simulation framework. Experimental results demonstrate that our co-designed tool morphology reduces real-world weighing errors by 45% relative to a standard tool, including on previously unseen materials. These results demonstrate our method can adapt both the control policy and the physical tool to the dynamics of the target material, bringing a new paradigm for material manipulation to the field of laboratory automation.

cs.RO↗

All-Microwave Multiqubit Gates

Implementing high-fidelity multiqubit gates is critical for reducing circuit depth in near-term quantum processors and fault-tolerant architectures. However, realizing multiqubit interactions and entanglement remains a critical challenge that limits the implementation of multiqubit gates. Here we propose an all-microwave scheme to realize single-step multiqubit gates incorporating n control and m target qubits, tailored for frequency-tunable superconducting transmon networks. By leveraging cross-resonance (CR) drives, this approach induces effective two-body ZX interactions that are significantly stronger than those achieved in conventional resonant regimes, circumventing the need for tunable couplers. The system's effective Hamiltonian is analytically derived using both quad frame rotation and rigorous block diagonalization. These combined theoretical methods facilitate the identification of optimal parameter regimes that simultaneously enhance desired target couplings and suppress parasitic higher-order terms, such as ZXX interactions. Through numerical simulations and gradient-based pulse-shape optimization, we demonstrate three-qubit gates achieving fidelities exceeding 99.9% within a 60 ns gate duration. Furthermore, we evaluate the fundamental scalability of this framework by extending the optimization to a five-qubit CXXXX architecture, confirming the physical viability of the multi-target driving scheme. Because these multiqubit operations naturally emulate stabilizer measurements, this architecture provides a hardware-efficient strategy for reducing the circuit depth of stabilizer-type operations by replacing sequential two-qubit gate decompositions with single-step multiqubit gates.

quant-ph↗

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Reconstructing dynamic driving scenes from recorded sensor data supports closed-loop evaluation of autonomous driving systems by synthesizing observations beyond the original trajectory. Unlike cameras and LiDAR, radar measures radial velocity directly through Doppler. Yet existing radar novel-view synthesis fails to exploit this capability: methods addressing dynamic scenes reconstruct only range-azimuth tensors, while methods that render Doppler assume static scenes. Moreover, because radar processing spreads each reflection across multiple bins, existing representations absorb this spread into scene geometry, causing it to render incorrectly when the viewpoint moves. We present DyRAD, which models dynamic driving scenes using static background reflectors and motion-tracked dynamic point reflectors to render complete range-azimuth-Doppler (RAD) tensors. Reflector velocities are derived from object tracks and projected onto the line of sight, making Doppler both a rendered output and supervision for those tracks. Crucially, we render reflectors through a fixed analytic point-spread function (PSF) derived from the radar's signal-processing chain, preventing sensor-induced spread from being baked into the scene representation. Beyond improving scene reconstruction, this separation also enables zero-shot sensor-configuration transfer, allowing the same reconstructed scene to be rendered under different radar specifications without refitting. We evaluate DyRAD on RADIal, Boreas, and a synthetic benchmark across both on-path poses and displaced viewpoints untested by prior work. On RADIal, DyRAD recovers radar detections in 90.7% of reference-detected objects, compared with 26.9% for the strongest baseline.

cs.CV↗

Grounding with Confidence: Controllable Generative Video Temporal Grounding

Video temporal grounding supports applications such as video search, content review, and automated editing by localizing events described in natural language. Yet existing generative models typically output timestamps without explicit interval-level confidence scores to guide candidate selection. We separate candidate generation from acceptance by scoring individual intervals within the original decoding pass. A lightweight confidence head reads pooled decoder states, providing an explicit score trained for interval selection. Offline verifier scores supervise the head on fixed candidate sequences, and temporal-overlap labels adapt it to current rollouts during reinforcement learning. GT-anchored candidate-pool supervision and set-level optimization train the generator. The resulting scores support ranking, threshold-based selection, and rejection without invoking an external verifier at inference. On a fixed OMTG-Bench candidate pool, confidence raises query-macro Recall@0.5 from 9.95% to 14.42% over generation order at a 10% global return budget, and from 26.48% to 31.12% at a 25% budget. The continuous scores let downstream applications adjust return budgets or acceptance thresholds to match their precision-recall preferences, without regenerating candidate intervals.

cs.CV↗

Markovian Dynamics Enforcer: Feasibility Preserving Correction on Learned Dynamics Manifolds

Neural trajectory predictors can reach low prediction error while violating dynamics, actuator limits, or state constraints, especially when controls are unobserved and dynamics are partially specified. We introduce the Markovian Dynamics Enforcer (MaDE), a time-invariant post-hoc operator mapping state-transition proposals onto a learned feasible dynamics manifold, trained on feasible states without ground-truth controls. For each transition it infers a control and recomputes the state through a completion model of known physics plus a learned residual. It then corrects that control by gradient-based inequality reduction, so inequality satisfaction is best-effort within an iteration budget. Since every correction iterate re-enters the completion model, the returned state is dynamically consistent by construction relative to that model and the supplied previous-state anchor. MaDE drives dynamics residuals to essentially zero on fully specified simulated systems, and on an underspecified system leaves a smaller true-dynamics residual than the baselines. Designed to attach to arbitrary predictors, the frozen operator is evaluated downstream of recurrent, structured state-space, and transformer predictors. On recorded vehicle trajectories the one-step residual against a kinematic bicycle model is 0.0071 to 0.0072 for MaDE and 0.1703 to 0.1714 for raw predictors. MaDE raises average displacement error by a factor of 1.57 to 1.83.

cs.LG↗

Zero modes and oscillatory instabilities of a Lorentz-violating Kalb-Ramond field on a Schwarzschild background

We study equilibrium configurations and linear perturbations of a Lorentz-violating Kalb--Ramond field with a quartic symmetry-breaking potential and a nonminimal Riemann coupling on a fixed Schwarzschild background. For static spherical configurations, the electric component is determined algebraically by a characteristic function that can develop a finite-radius double root. Approaching this degenerate configuration, the monopole electric response scales as $|\widehat e(ω,r_c)|\propto(γ_c-γ)^{-1/2}$, while the propagating monopole amplitude remains regular, showing that the enhancement originates from the algebraic constraint rather than from a dynamical instability. For higher multipoles, we obtain an exact tower of zero-frequency modes, $ξ_{\ell n}=-(\ell+n+1)(\ell+n+2)/3$. For $\ell=1$, the finite-frequency spectrum contains two distinct low-frequency branches. As the Riemann coupling becomes more negative, the corresponding purely imaginary unstable modes coalesce and leave the imaginary axis as $ω_\pm=\pmω_R+iω_I$, producing an oscillatory instability. Near the merger, the real-part splitting follows a square-root law.

gr-qc↗

AIMS: An Agentic AI Framework for Sim-to-Real Multi-Modal ISAC

Multi-modal integrated sensing and communication (ISAC) enables environmental perception and reliable connectivity for intelligent wireless networks. Data-driven multi-modal ISAC models depend heavily on annotated real-world data to learn relationships across sensing and wireless observations, thereby constraining scalable deployment. Although synthetic data generation reduces the burden, adapting existing simulation pipelines to a target deployment requires consistent scene, sensing, wireless, and learning configurations, while mismatches among these coupled components impair sim-to-real transferability. To address the challenge, we propose an agentic artificial intelligence (AI) framework for sim-to-real multi-modal ISAC, named AIMS. Given a natural-language deployment request specifying the target task, deployment conditions, and real-data budget, AIMS derives a deployment-specific sim-to-real configuration and coordinates its execution to produce a deployment-specific task model. A two-agent architecture coordinates scene construction with task learning. A scene construction agent generates geographically grounded, synchronized sensing and wireless records from shared physical states, while a scene understanding agent configures task-relevant modalities and mixture-of-experts (MoE) learning for zero-shot inference or few-shot adaptation. Structured domain knowledge guides dependency-aware planning, while validation evidence supports feedback-driven revision of affected decisions. Experiments on the real-world DeepSense 6G dataset demonstrate improved vehicle detection and beam prediction over the considered simulation and fusion baselines. A separate orchestration benchmark evaluates task interpretation, dependency reasoning, and feedback-driven replanning across diverse deployment requests, showing improved plan correctness with structured domain knowledge and validation feedback.

cs.AI↗

Single-shot state preparation threshold: rigorous theorem and statistical mechanical mapping

Preparing logical codewords with a single round of noisy syndrome measurements can reduce the time overhead of fault-tolerant quantum computation. It remains unclear which structural properties of the code suffice to guarantee a threshold for single-shot state preparation. Here we prove that linear confinement suffices for a nonzero single-shot state-preparation threshold for CSS quantum low-density parity-check codes, in contrast to the common belief that soundness is necessary. When the distance is at least logarithmic, we prove that below nonzero local stochastic readout and data error thresholds, the logical error probability after a formal ideal final recovery is exponentially small and approaches zero when the code size approaches infinity. Under the stronger assumption of linear soundness, we show that the residual data error of the preparation protocol is local stochastic, directly allowing composition with other fault-tolerant gadgets. To investigate single-shot state-preparation threshold beyond the above assumptions, we further develop a statistical mechanical mapping that expresses the final logical error rate in terms of a classical partition function. We perform Monte Carlo simulations of this model for the three-dimensional toric code. The numerical results suggest that sublinear soundness can also support a single-shot state-preparation threshold.

quant-ph↗

Strong Dimerization and Field-Induced Reconstruction of the Low-Energy Spectrum in $\mathrm{Cu}_3(\mathrm{OH})_4(\mathrm{HCO}_2)_2$

We investigate the field-dependent low-energy thermodynamics of the distorted triangular quantum antiferromagnet $\mathrm{Cu}_3(\mathrm{OH})_4(\mathrm{HCO}_2)_2$ using a sector-resolved Superblock Diagonalization Method (SBDM) supplemented by a transfer-matrix treatment of weakly coupled layers. The strong exchange hierarchy, dominated by the Cu1--Cu1 intradimer coupling $J_2=150\,\mathrm{K}$, separates a high-energy dimer sector from a much softer magnetic manifold formed predominantly by the Cu2 moments. In the 24-site cluster, sixteen Cu1 spins form the strongly bound sector while eight Cu2 spins remain magnetically active; polarization of these eight spins gives $S^z=4$ compared with $S^z_{\rm sat}=12$, providing a direct microscopic origin for the one-third magnetization scale. The finite-temperature thermodynamics reveals a non-monotonic field evolution of the low-energy scale: the dominant $C/T$ feature softens with increasing field, reaches a minimum near the field region around $2\,\mathrm{T}$, and subsequently hardens as the low-temperature magnetization approaches $M_{\rm sat}/3$. Temperature and field sweeps thus expose a common field-induced spectral reconstruction, while the strongly reduced entropy reflects the restricted number of thermally active degrees of freedom below the dimer excitation scale. Our results identify strong-dimer-induced reduction of the active magnetic Hilbert space, followed by field-driven reorganization of the residual spin sector, as the common microscopic origin of the one-third magnetic response and the non-monotonic low-temperature thermodynamics.

cond-mat.str-el↗

Learning to Explain While Planning: Rule-Aligned Diffusion Planning for Autonomous Driving

Diffusion planners exhibit strong capabilities in generating multimodal trajectories. However, existing methods primarily rely on expert demonstrations to fit trajectory distributions, learning statistical correlations among scenes, behaviors, and trajectories without explicitly modeling driving rules. In long-tail scenarios where expert data are scarce, the lack of behaviors to imitate may lead to trajectories that violate safety or compliance requirements. Moreover, their generation process lacks rule-level explanations, making it difficult to determine which rules drive trajectory adjustments, when they take effect, and how strongly they act, thereby limiting failure diagnosis, safety validation, and targeted improvement. To address these limitations, we propose the Rule-Aligned Diffusion Planner (RADP), which incorporates differentiable driving rules into the diffusion objective during training, turning rule knowledge into intrinsic behavioral principles beyond finite demonstrations. We further introduce Rule-Pressure Attribution (RPA), which constructs supervision signals from gradients of rule losses with respect to predicted trajectories and employs a lightweight attribution head to estimate the optimization pressure exerted by each rule online. To assess the closed-loop behavioral relevance of these attributions, we propose a temporal risk-alignment protocol that evaluates whether current rule pressures reflect corresponding risks during subsequent closed-loop execution. Experiments on nuPlan show that RADP improves closed-loop planning in challenging safety-critical scenarios, while RPA exhibits consistent temporal alignment with subsequent rule-specific risks, validating both intrinsic rule learning and rule-level interpretability.

cs.LG↗

Fenchel Tilting: Weighted Correction for Efficient Finetuning of Generative Models

Adapting a pretrained generative model to an arbitrary preference expressed as a utility function underlies reward alignment, guided design, and constraint satisfaction, enabling diverse applications. Existing fine-tuning methods trade off generality against computational cost: they either restrict the family class of supported preferences to keep optimization simple or preserve generality at the expense of efficiency. We introduce Fenchel Tilt Flow Control (FTFC), which decouples utility optimization from generative-model fitting. FTFC first optimizes for a target distribution by jointly fitting an effective reward and density-ratio weights on pretrained samples. Method combines the utility's variational structure with Fenchel duality, supporting general $f$-divergence penalties that determine how rewards are transformed into an distribution-correction weights. These weights are then frozen and used to modify a diffusion or flow model in a single stage of importance-weighted denoising or flow matching, without differentiating through sampling trajectories. We establish exact duality for concave utilities under suitable conditions and show that weighted fitting reproduces the optimal target distribution for a given utility. Across image and molecule generation benchmarks, FTFC improves over baselines on diverse preference functions, while also being up to $20\times$ more efficient. roposed method enables adaptation beyond expected-reward maximization without complex optimization, while preserving robustness for more general class of the utility functions compared to baselines.

cs.LG↗

Probing Collective and Individual Kondo Screening: Multi-Stage, Multi-Channel Kondo Effects in a $C_3$-Symmetric Four-Impurity Model

Collective screening has been proposed to underlie the basic physics of multi-impurity and lattice Kondo systems. But how to establish this picture and distinguish it from individual (local) Kondo screening remains a grand challenge. The recently developed auxiliary-bath numerical renormalization group (AUNRG) method provides a key step towards resolving this issue. Its application to the $C_3$-symmetric three-impurity Kondo (3IK) model with a shared electron bath reveals fully screened Fermi liquid ground states arising from collective screening of cluster spin degrees of freedom at small Kondo coupling $J_{\rm K}$ and individual (local) screening of local impurities at large $J_{\rm K}$. Here we design a four-impurity Kondo (4IK) model where the probe impurity couples only to the original three Kondo impurities to detect the nature of the screened states. We show that the collective and individual Kondo screenings give rise to emergent multi-stage, two-channel Kondo effect and unstable three-channel Kondo effect, respectively. This suggests a special helicity structure of the screened states and confirms the idea of collective screening in multi-impurity Kondo systems. Our work also demonstrates that the same auxiliary-bath construction can be readily extended to models with additional impurities to explore novel many-body quantum states with emergent cluster degrees of freedom.

cond-mat.str-el↗

Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training

An unsuccessful LLM agent rollout contains more information than its final reward: the observations available to the agent, the actions it chose, and the environment's responses. Reusing this experience for learning requires identifying a decision to revise and testing a concrete alternative. We introduce the Agent Error Dataset (AED), comprising 50,228 error-diagnosis pairs from 9,961 source tasks across 33 environments, 19 harness families, and 23 policy models in text-based agent systems. We retain source traces and execution metadata to support cross-setting failure analysis and re-diagnosis without repeating the original rollout. Our five-stage Agentic Error-to-Training (AET) pipeline collects natural failures, generates diagnoses and proposed corrections, and checks them against recorded evidence. Where replay is supported, we compare corrections with original-action retries from the same checkpoint under matched execution settings. We then construct separate training views for diagnosis and actor recovery. Across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1%, a gain of 32.7 percentage points. Using a separately frozen diagnosis release, full-diagnosis fine-tuning on 1,656 source tasks raises Qwen3-8B's exact-step agreement with internal teacher labels from 47.2% to 63.6%, averaged over three seeds on a 943-case holdout. The strongest prompted reference in this comparison scores 54.7%, and mean agreement improves at each of four increasing training-set sizes. In a single-seed comparison of actor-training recipes, action-only repair training scores 6.67 percentage points higher on WebShop-lite than success-only training.

cs.AI↗

Less is more: error-distance scaling relation for data-efficient kilometer-scale downscaling of extreme heat

Extreme heat is where urban adaptation needs kilometer-scale data the most, but the simulations training a downscaler can cost more than they save, and how much is needed has not been identified. We measured it with CASPER, a U-Net with a structure-preserving loss downscaling 32 km reanalysis to 1 km temperature, humidity and wind, across 24 configurations of one to eight months. Held-out error grows linearly with climatological distance to the training data, RMSE = 0.83 + 2.95 d, explaining 90% of its variance against 7% for volume and predicting unseen months in advance. On held-out extreme summer weeks CASPER preserves the fine-scale structure and cross-variable physics that matched-budget baselines degrade, and matches station observations during documented heat waves to within 1.8 K. Transfer to a new region degrades geographically; 11 days of local simulation cuts Vancouver's held-out error from 3.8 to 1.3 K. Training periods should span the target climate: the same accuracy for four times less simulation, putting kilometer-scale downscaling of extreme heat within reach of groups without large computing facilities.

physics.ao-ph↗

Social-WM: Safety-Aware Latent World Models for Robot Social Navigation

Safe social navigation requires a robot to anticipate not only the future consequences of its actions, but also whether a nominal action can actually be executed under surrounding physical and social constraints. We present Social-WM, an efficient latent world-model planning framework trained from egocentric RGB video sequences. Our key observation is that social-navigation experience contains a systematic discrepancy between the nominal action and the realizable action: a nominal forward action may be fully executed in free space, but needs to be constrained when heading towards a pedestrian or obstacle. Social-WM learns these safety-relevant consequences directly through action-conditioned future prediction, where the target is the actual observed future following each command. We further introduce a realizable inverse-dynamics objective that associates observed latent transitions with the action actually realized rather than the nominal one. At deployment, candidate actions are imagined through the latent world model, and the inverse dynamics model estimates their realizability; nominal--realizable discrepancy then provides a safety signal before execution. The learned dynamics and realizability model remain goal-independent and support both position- and image-goal navigation. On Social-HM3D, Social-WM achieves 63.77% success while reducing human collisions to 21.67%, and maintains strong performance under zero-shot transfer to Social-MP3D, without explicit pedestrian tracking, privileged human state, or online reinforcement learning.

cs.RO↗

Quantum Markov State Models for Metastable Dynamics

Open quantum systems can rapidly lose most microscopic information, leaving only a few degrees of freedom to govern their long-time dynamics. Classical Markov state models (MSMs) describe such metastable dynamics as transitions among a few representative phases and are widely used to reduce complex-system dynamics in condensed matter and chemical physics. In quantum systems, however, phase labels alone are insufficient when coherence persists between metastable states. Even when the slow modes are known, their spectral projection need not produce valid quantum states. We construct quantum Markov state models (QMSMs) that describe the surviving information and its evolution on a small physical state space containing classical sectors and quantum matrix blocks. We quantify metastability by assuming that the evolution channel $\mathcal C$ changes little when applied a second time, with sufficiently small defect $η=\|\mathcal C^2-\mathcal C\|_\diamond$, and that the number of slow degrees of freedom is bounded independently of the full system size. Under these assumptions, we construct compression and reconstruction channels whose composition recovers every reduced state exactly, while the reverse composition gives an exactly idempotent channel approximating $\mathcal C$, answering Kitaev's exact-rounding question [Kit25]. Our constructions give an optimal diamond-norm bound of $\mathcal{O}(η^{1/3})$ on all input states, as well as an improved bound of $\mathcal{O}(η^{1/2})$ on metastable states prepared by $\mathcal C$, with constants depending only on the slow dimension. The reduced transition channel can be iterated to predict the microscopic dynamics with controlled error. We illustrate the QMSM through a weakly driven dissipative spin chain supporting either a metastable logical qubit or long-lived classical phases.

quant-ph↗

STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction

Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding, capabilities that align with the nuanced requirements of social robot navigation. However, it remains unclear whether VLMs can accurately understand complex social navigation scenes (e.g., inferring the spatial-temporal relations among agents and human intentions), which is essential for safe and socially compliant robot navigation. While some recent works have explored the use of VLMs in social robot navigation, no existing work systematically evaluates their ability to meet these necessary conditions. In this paper, we introduce the Social Navigation Scene Understanding Benchmark (SocialNav-SUB), a Visual Question Answering (VQA) dataset and benchmark designed to evaluate VLMs for scene understanding in real-world social robot navigation scenarios. SocialNav-SUB provides a unified framework for evaluating VLMs against human and rule-based baselines across VQA tasks requiring spatial, spatiotemporal, and social reasoning in social robot navigation. Through experiments with state-of-the-art VLMs, we find that while the best-performing VLM achieves an encouraging probability of agreeing with human answers, it still underperforms simpler rule-based approach and human consensus baselines, indicating critical gaps in social scene understanding of current VLMs. Our benchmark sets the stage for further research on foundation models for social robot navigation, offering a framework to explore how VLMs can be tailored to meet real-world social robot navigation needs. An overview of this paper along with the code and data can be found at https://larg.github.io/stars.

cs.RO↗

ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents

Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privileged rescoring, but directly applying it to CUA online training presents two challenges: fixed guidance may become misaligned with the student's current state, and guidance-induced probability shifts may conflict with step-level correctness. We introduce ComputerSD, an online self-distillation method for CUAs that converts real-time feedback from executed GUI transitions into guidance for policy learning. A fine-tuned GUI analyzer produces guidance and a step-level value score after each action; the guidance provides privileged context, while the score regulates the resulting OPSD signals. ComputerSD jointly optimizes token-level OPSD and trajectory-level GRPO in a fully asynchronous training framework. On OSWorld-Verified, ComputerSD outperforms outcome-only GRPO by 1.9 and 4.1 percentage points on the general-purpose Qwen3-VL-8B-Thinking and specialized EvoCUA-8B backbones, respectively. Evaluation in out-of-distribution settings further supports the generalizability of ComputerSD. These results demonstrate the effectiveness of learning from real-time feedback through online self-distillation for CUAs.

cs.CV↗