Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 559 records · Page 31Linked to original sources

EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data

Web agents often revisit the same sites, yet most evaluations discard the procedures learned in earlier successful interactions. We introduce EconSkills, a skill library and evaluation framework that distills verified EconWebArena trajectories into parameterized standard operating procedures for retrieving live economic data. Each skill records its scope, navigation procedure, site-specific guidance, verification checks, and recovery steps while replacing source-instance values with placeholders. EconSkills separates two questions: whether a known relevant procedure transfers to a held-out task, and whether an agent can retain that benefit when selecting from a library. In controlled transfer, matched skills improve success over no-skill prompting and require fewer steps on paired successes. Under the evaluated prompt formats, the parameterized skill prompt substantially outperforms the corresponding raw-trajectory prompt. With the full 50-skill library, retrieval is competitive with the no-skill baseline overall and performs best on tasks with a direct family match; approximate matches on other tasks offset these gains. Browser trajectories further identify when procedural guidance shortens portal-specific navigation and when semantic verification remains necessary. These results establish that reusable economic web procedures can transfer across task instances and provide a concrete design target for match-aware selection and context delivery.

cs.AI↗

MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration

The deployment of Vision-Language Models (VLMs) on edge devices is severely bottlenecked by memory bandwidth, necessitating aggressive sub-8-bit quantization. Since edge accelerators are strictly constrained by area and power, they require end-to-end quantized models. However, the extreme dynamic range gap between multi-modal tokens causes standard block formats to suffer "microscaling collapse," where a single massive outlier hijacks the shared exponent, underflowing surrounding elements and destroying attention maps. To break this bottleneck, we propose Micro-Inverted-Scaling (MiX), a novel format that mathematically inverts the microscaling paradigm: rather than grouping multiple mantissas under one shared exponent, MiX groups private, per-element exponents under a single shared mantissa. To handle asymmetric VLM outlier topologies, we introduce an adaptive dual-format (MiX-MX) inference framework. By algebraically factoring out the shared MiX mantissa, this framework maps to a custom accelerator, replacing multipliers with efficient shifters. Evaluated end-to-end on multiple VLMs, our 4.5-bit MiX formulation exhibits equivalent or superior accuracy on multi-modal benchmarks compared to NVFP4. Simultaneously, the MiX accelerator delivers a 25% improvement in area efficiency over the NVFP4 baseline and a 2.3-4.5x speedup with 1.4-2.9x energy reduction across models compared to the state-of-the-art accelerator Focus, proving the inverted-scaling datapath is physically superior for efficient VLM deployment.

cs.AR↗

Non-Thermal Effects in Fermionic Atoms Coupled to Open Cavities

We study a Fermi-Hubbard model coupled to a open dissipative single cavity mode. Using Keldysh diagrammatics we derive and solve the quantum kinetic equations for the fermions, taking the bosonic cavity mode as a source of non-equilibrium noise and dissipation. In absence of Hubbard interactions we show that the fermions reach generically a non-equilibrium steady-state, characterized by a non-thermal distribution function. Quite interestingly we demonstrate that the latter exactly nullifies the heat-current between fermions and cavity mode. We discuss the regimes of parameters where a low-frequency effective temperature description emerges and how the fluctuations affect the mean-field phase diagram for the superradiance phase transition. Finally we include Hubbard interaction in the weak-coupling regime and show that it leads to a crossover towards a full equilibrium distribution.

cond-mat.str-el↗

HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface

Large-scale vision-language-action (VLA) models provide powerful priors for robot manipulation, yet adapting them to a specific deployment remains challenging. Supervised fine-tuning (SFT) on task-specific demonstrations provides a step toward deployment, but faces two persistent limitations: static data provide limited coverage of out-of-distribution states, and standard imitation objectives do not distinguish progressing behavior from less useful data. Interactive post-training can address these limitations, but typically requires repeated policy execution and human intervention on a physical robot. We introduce HIL-UMI, a policy-guided Universal Manipulation Interface (UMI) framework for robot-free human-in-the-loop VLA post-training. During handheld UMI demonstrations, HIL-UMI queries the current policy on the same observation stream without executing its predictions. The Energy Score compares the human action trajectory with policy inference and triggers collection when their discrepancy indicates an out-of-distribution region. In a separate feedback loop, low online advantage predictions identify essential segments for refining a progress-based advantage estimator. The updated estimator then guides advantage-conditioned behavioral cloning using a balanced mixture of base demonstrations and new policy data. This design preserves the iterative and policy-aware nature of human-in-the-loop learning while decoupling data collection from robot deployment. Experiments on four real-world tasks spanning long-horizon and precise manipulation show that HIL-UMI achieves consistent improvement over SFT and benefits from both targeted collection and advantage refinement. Moreover, HIL-UMI outperforms HG-DAgger on Clean Up Table with lower per-frame collection time, suggesting a scalable path for VLA post-training across operators and locations.

cs.RO↗

A comparison between statistics and 't Hooft anomaly

Generalized symmetries and topological excitations, as well as symmetry anomalies and the statistics of topological excitations, are widely believed to be related. There are, however, pitfalls in how this relation is established. A lattice truncation of a symmetry transformation to a finite patch gives a symmetry patch operator that creates symmetry defects at its boundary. This geometric picture resembles a hopping operator creating topological excitations at the boundary of its support, but does not provide well-defined statistics, let alone guarantee agreement with the symmetry anomaly. A more natural and robust relation is that the hopping operators of topological excitations are symmetric: they commute with symmetry transformations. Under suitable assumptions, this condition yields a one-to-one correspondence between statistics and anomalies. We further couple boundary matter to a DW gauge field in one higher dimension to explain this relation from the perspective of gauging. A hopping operator is, in essence, a gauge-invariant operator acting on the physical degrees of freedom after gauging. Once the gauge-field background is fixed, global symmetry comes from gauge transformations that preserve that background, while the symmetric condition on hopping is precisely the remaining gauge invariance. This distinction clarifies potential misconceptions in the literature and provides a more reliable framework for comparing symmetries and topological excitations.

cond-mat.str-el↗

Embedding Models Measure in Peculiar Ways

Embedding spaces define notions of semantic similarity and distance. We study whether those embeddings reflect physical measurements of mass, distance, time and volume, which admit a unique, objective notion of semantic equivalence and distance. We find that physical measurement is only weakly modeled in the embedding space, and that instead quite peculiar measurement patterns can be observed. Further analysis indicates that embedding representations of physical measurements are strongly influenced by superficial string similarity, and recalibration of similarity does not substantially improve the alignment.

cs.CL↗

Generation of Stable Peak-Power Similaritons through Gain-Managed Nonlinearity

Fiber lasers and amplifiers offer attractive alternatives to conventional solid-state systems. However, generation of high-energy ultrashort laser pulses in fibers faces challenges due to the complex interplay of multiple nonlinear effects arising due to pulse confinement within a small fiber core and also limitations imposed by the gain bandwidth of the available active fibers. The discovery of self-similar amplification and gain-managed nonlinear amplification (GMNA) pulse propagation regimes in fibers with normal dispersion suggests that these challenges can be turned into an advantage. Here we show that pulses generated in the GMNA regime are, in fact, the realization of the idealized similariton-type pulses in realistic fibers with limited gain bandwidth. Our analytical and numerical results show how one should shape the fiber gain as a function of propagation length to achieve constant peak power similariton-like pulses with steadily increasing energy, the pulse bandwidth exceeding the gain bandwidth, and the nearly linear frequency chirp allowing for efficient pulse compression to its Fourier limit. Absent Raman nonlinearities, these pulses can reach $μ$J level energies in standard single-mode fibers, representing a tenfold increase in pulse energy compared to the best currently available nonlinear amplifiers. Our results have significant implications for the fundamental understanding of nonlinear wave dynamics and for the advancement of fiber laser technology, supporting the reliable generation of high-energy pulses for practical use in areas such as micromachining, metrology, and bioimaging.

physics.optics↗

Energy stability and error estimates for a second-order structure-preserving exponential integrator method for smectic-A liquid crystals

In this work, we develop a second-order, linear, decoupled, and structure-preserving numerical scheme for the modified Landau--de Gennes model of smectic-A (SmA) liquid crystals. The main contributions are threefold. First, to the best of our knowledge, we propose the first integration of the generalized scalar auxiliary variable (GSAV) approach with a second-order exponential time-differencing Runge--Kutta (ETDRK2) discretization, leading to a second-order GSAV--ETD2 scheme. Second, we prove that the proposed scheme satisfies an unconditional energy-dissipation law, thereby closing the theoretical gap in the energy-stability analysis of second-order GSAV exponential integrators of this class. Third, by deriving a coercive discrete reformulation, we establish a fully discrete error estimate without imposing any coupling condition between $τ$ and $h$, achieving the optimal convergence rate $\mathcal{O}(τ^2+h^2)$. Numerical experiments are presented to verify our theoretical results and to simulate the self-assembly dynamics of the SmA phase.

math.NA↗

Why most neutron star low-mass X-ray binaries accrete transiently: an evolutionary study of transient and persistent phases

A neutron star (NS) low-mass X-ray binary (LMXB), in which an NS accretes matter from a low-mass donor star, is an ideal source for probing some fundamental aspects of physics and astronomy, such as strong gravity, superdense matter, and the accretion-ejection processes. However, to reliably achieve these goals, one must adequately understand NS LMXBs, including why some accrete persistently and others transiently. Focused models, such as those based on a thermal-viscous instability in the accretion disk, are considered to explain transient accretion. However, broader perspectives, including which LMXB parameter values and phases cause transients and why there are more transients than persistents, remain poorly understood. Here, our computation of the long-term evolution of NS LMXBs addresses these questions, providing insight into LMXB parameters and phases, naturally producing more transients than persistents, and being partially consistent with the known properties of observed sources. For example, we typically find a greater fraction of persistent phase at lower orbital periods from the LMXB evolution computation, which is somewhat consistent with observations. However, a lack of full consistency calls for improving the aforementioned focused models, and our computations provide a new way to discriminate among these models.

astro-ph.HE↗

DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

Online agent deployments accumulate execution trajectories at massive scale and behavioral diversity, for which predefined annotation criteria hardly exist. Extracting useful evidence therefore demands costly manual annotation or verifier signals that fail to scale, leaving valuable evidence buried among redundant, incomplete, and failed executions. This raises a question: without post-execution rewards or correctness labels, how can reusable experience be distilled from the trajectories themselves? To address this challenge, we introduce DENSE (Distilling Evidence from Nested Subtask Executions), which organizes trajectory-derived evidence into nested shortcut trees. By consolidating redundant attempts, identifying resolved subtasks, and retaining useful steps alongside outstanding requirements, DENSE transforms noisy execution traces into structured and reusable task-solving feedback. To evaluate whether such feedback helps agents retry the same task, we design REFIT, which measures success-rate changes between the initial attempt and feedback-guided retries. Among feedback methods without external outcome supervision, DENSE achieves the highest strict pass rate across four agent models on Terminal-Bench 2.1, improving over initial attempts by 7.12-21.81 percentage points with 19.0-43.6% fewer agent tokens on retries. In addition, on hard tasks DENSE consistently outperforms self-reflection in cumulative pass rate across multiple feedback iterations on all four models, demonstrating its strong potential for continual agent self-improvement.

cs.AI↗

Signature of mechanically induced cell extrusions in cell size distribution

How a growing tissue organizes its own homeostatic state is a central question in the physics of living matter. We show that when a growing epithelial sheet counteracts increasing cell density by mechanically squeezing cells out of its plane, a homeostatic in-plane pressure emerges as a generalization of a yield stress. We find that in the quasistatic growth limit the homeostatic state is marginally stable, with a pseudogap in the distribution of local distances to the extrusion threshold pressure. Because such mechanically induced extrusions arise from an instability of individual cells, the pseudogap is imprinted in the distribution of cell areas. This provides an image-based way to test for presence of mechanically induced extrusions and we identify this signature in the developing wing epithelium of \textit{D.~melanogaster}. We expect the same principles to apply to confined three-dimensional tissues.

physics.bio-ph↗

A Second-Order Maximum-Bound-Preserving and Energy-Stable Exponential Time-Differencing Method for Allen--Cahn-Type Gradient Flows

The energy dissipation law and the maximum bound principle (MBP) are two important physical features of the well-known Allen--Cahn equation. In this paper, we develop and analyze novel second-order linear numerical schemes for a class of Allen--Cahn type gradient flows. Our scheme is based on the generalized scalar auxiliary variable (GSAV) approach and a novel second-order exponential time-differencing Runge--Kutta (ETDRK2) method. The resulting formulation overcomes a longstanding difficulty in combining these two techniques while retaining both the MBP and energy stability. We prove that the proposed scheme unconditionally preserves both the MBP and energy stability. In addition, rigorous error analysis is carried out for the proposed scheme, establishing second-order accuracy in both time and space without imposing any coupling condition between the time step $τ$ and the spatial mesh size $h$. We also present some numerical experiments to demonstrate the efficiency of the proposed scheme and its preservation of the theoretical properties.

math.NA↗

SparkDiffusion: Mitigating the High-Sparsity Trap --- A Unified Framework for up to $265\times$ Single-GPU Acceleration of Visual Generation

Video diffusion transformers are expensive because attention dominates long spatiotemporal token sequences. We identify the \emph{high-sparsity trap}: at extreme attention sparsity, step-local training losses keep decreasing while terminal generation quality stagnates or degrades. The trap is one of supervision: the dominant terminal errors originate in the high-noise structure-generation stage, and terminal-aligned training corrects terminal errors that substantially extended step-local training cannot. This yields a simple staging principle: \emph{first adapt the sparse architecture into a coarse prior, then correct the terminal distribution}. We instantiate the principle as \method, a unified acceleration framework for visual generation that combines a short sparse warm-up, few-step trajectory-mixed distillation, and FP8 quantization with fused kernels. \method sustains $97\%$ attention sparsity with strong visual quality on long-sequence 720P generation across Wan2.1/Wan2.2 backbones and T2V/I2V tasks, and $90\%$ sparsity on Wan2.1-T2V-1.3B-480P. With 3-step CFG-free inference, \method achieves a $265\times$ end-to-end speedup over the 50-step CFG dense baseline for Wan2.1-T2V-14B-720P on a single RTX~5090 ($220\times$ on H100), and denoises a Wan2.1-T2V-1.3B-480P video in $1.3$s.

cs.CV↗

CTRL: Control-Based Time Series Forecasting with LLM-Guided Residual Learning

Time series forecasting underpins critical decision-making across diverse domains. While large language models (LLMs) offer promising reasoning capabilities, existing LLM-based time series forecasting approaches either reduce them to numerical predictors that bypass their strengths, or allow direct forecast generation that destabilizes predictions in non-stationary settings. We introduce CTRL, a framework that decouples semantic reasoning from quantitative prediction. A frozen backbone generates base forecasts, while specialized LLM agents function as controllers that analyze backbone prediction errors through decomposed trend, seasonal, and irregular components, grounding reasoning in interpretable temporal structure. Each agent outputs compact control signals that a lightweight residual decoder translates into forecast corrections. CTRL incorporates label-free test-time adaptation that detects distribution shift from input statistics alone and readapts control signals with only 3-24 LLM calls via caching. CTRL is explicitly designed to improve robustness under non-stationary temporal dynamics and distribution shift, while remaining competitive on highly stationary time series where adaptive correction provides limited additional benefit.

cs.LG↗

RoboTalk: Learning Multi-Robot Communication and Coordination from Multimodal Demonstrations

Multi-robot collaboration could enable more efficient and scalable solutions to complex robotic tasks, but collaboration under partial observability remains challenging. Natural-language communication offers a promising approach to coordinating robots under partial observability. However, in decentralized manipulation, jointly learning explicit inter-robot communication and skill-level action selection from multimodal demonstrations remains underexplored for small vision-language models (VLMs) intended for on-device deployment. To address this gap, we introduce RoboTalk, a synthetic data-generation pipeline and dataset of 7,950 multimodal trajectories spanning 53 mobile-manipulation kitchen tasks for training small VLMs to communicate and coordinate. The dataset includes a leader-follower planning protocol, tool calls (perception, manipulation, navigation, and communication), rationale traces, and diversified natural-language communication. Fine-tuning open-source models on our dataset can reach 77% success on novel held-out tasks, a significant improvement over the untuned open source models, which had a success rate of around ~2%.

cs.RO↗

MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents

The effectiveness of agent memory ultimately depends on whether the underlying LLM gives each memory in context an appropriate degree of influence over its response. Yet this capability has remained largely overlooked. To assess this capability, we introduce MemCalib, a benchmark grounded in realistic memory-system scenarios for evaluating memory use and advancing optimization algorithms. Results on the MemCalib test set reveal that frontier open- and closed-source models struggle to use memory appropriately. They frequently over-use or under-use memory rather than matching each proposition's actual use to its target level, leading to biased, low-quality responses. Experiments with common post-training algorithms, including group relative policy optimization and on-policy self-distillation, further reveal a clear directional skew: trained models improve in one direction while deteriorating in the other. We therefore propose MemCalib-RL, an ordered bidirectional counterfactual credit-assignment algorithm that separates over- and under-use signals and localizes their credit to response tokens through exact atom ablation. Results across model families and scales (Qwen3-8B, Ministral-3-8B-Instruct, and Qwen3.5-35B-A3B) show that MemCalib-RL achieves the best overall performance while better balancing over-use and under-use, with gains generalizing beyond MemCalib in external benchmark evaluation. Further experiments support its design choices and robustness and provide insight into its training dynamics.

cs.LG↗

A decision procedure for intuitionistic modal logic IS4 (and IK4)

In this paper, we show that the two intuitionistic modal logics IS4 and IK4 are decidable. We provide a constructive decision procedure, that, given a formula, produces either a proof showing the formula to be valid or a finite countermodel falsifying the formula, thus also proving the finite model property for both logics. The main ingredient of our strategy is the introduction of (possibly unsound) loop rules, which encode repeating behaviour in proof search. This paper fixes a previous mistake in our LICS'23 contribution.

cs.LO↗

Localized Decoherence as a Constructive Tool: An Optical Simulation of Matter-Wave Interference

Decoherence is conventionally regarded as a detrimental process that suppresses quantum coherence and limits the performance of quantum technologies. Contrary to this view, we show that spatially localized decoherence can instead be exploited as an interferometric tool. Beyond its conceptual implications, localized decoherence provides a novel approach to prepare arbitrary macroscopic spatial superpositions for testing quantum physics with massive particles. We demonstrate the proof of principle by making use of the correspondence between optics and matter-waves through the equivalence of the respective propagation functions. By employing stochastic phase fields to the optical counterpart to model the decoherence, our experiments confirm the expected interference. The results establish localized decoherence as a resource for quantum state engineering.

quant-ph↗