Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Ground-to-Cable Strain Transfer in Unburied DAS on Earth and the Moon

Distributed acoustic sensing (DAS) cables are typically buried to ensure good coupling and reduce unwanted atmospheric noise. Unburied deployments are attractive for rapid response applications and for missions to the Moon, where burial is impractical. Unburied DAS, however, suffers from degraded signal quality due to poor strain transfer. Here we quantify bending stress relief as a mechanism behind this loss: cable segments suspended between ground contact points accommodate ground strain by bending, so less axial strain reaches the fiber. We derive the strain transfer of such a segment analytically, compare it to numerical models, and validate it on 56 laboratory configurations of different cable types and lengths. The efficiency is governed by a single dimensionless parameter $Θ$, set by the ratio of the segment's sag to the cable radius. Keeping the loss below 3% requires a sag below about a quarter of the cable radius. The prediction is quasi-static, holding below the fundamental resonance of the segments. The sag is dominated by the curvature the cable retains from spooling and handling rather than by gravitational sagging, so $Θ$ has to be measured and not computed from cable specifications. Because the gravitational contribution is small, the criterion applies equally to deployment on the Moon. Lunar gravity does reduce the friction available to resist slip at the contact points, but a force balance indicates that slip remains unlikely for natural moonquakes.

physics.geo-ph↗

Simplex Relaxation for Discrete Diffusion

Discrete diffusion models for categorical generation are defined by a corruption kernel, which determines the intermediate state space and the associated reverse prediction problem. Within this family, uniform diffusion has been extensively developed, with recent work connecting its categorical corruption process to continuous representations and dynamics. Motivated by this view, we ask whether uniform discrete diffusion can be augmented with an explicit continuous state while leaving its categorical corruption process unchanged. We introduce Simplax, an exact Dirichlet--categorical augmentation that couples each corrupted categorical state with an auxiliary simplex-valued variable while preserving the uniform diffusion process as its categorical marginal. This augmentation yields a tractable Rao--Blackwellized reverse-bridge objective and a stochastic reverse sampler, while retaining the corrupted categorical state as the denoiser input. Empirically, Simplax improves the generative perplexity--entropy tradeoff on unconditional OpenWebText generation. On Sudoku, a model trained exclusively on $30$-clue puzzles achieves the highest accuracy among the compared methods across all evaluated clue densities, including the minimum uniquely solvable $17$-clue regime, and also achieves the highest validity in unconditional generation.

cs.CL↗

Large Language Model-Driven Small-Capitalization Trading: Integrating Financial News Sentiment, Macroeconomic Indicators, and Technical Signals

Large language models can extract richer signals from financial news than fixed sentiment lexicons, and recent work has explored feeding such signals into portfolio construction. We study an uncertainty-aware construction that feeds model-predicted risk -- decomposed into aleatoric and epistemic components -- directly into the covariance matrix of portfolio allocators, rather than treating portfolio risk as fixed or adjusting only expected returns. We evaluate the pipeline on Russell 2000 equities under three stock-selection regimes: a pure-alpha trigger that isolates abnormal stock moves not explained by macro indicators, a pure-beta trigger that captures macro-indicator moves before the stock itself fires, and a beta trigger in which both channels agree. Across the full holding-period grid, the separated pure-alpha and pure-beta legs usually dominate the beta intersection on Sharpe and return. Two horizons are especially informative. At one day, pure beta can work under low and moderate transaction costs because it captures immediate lead-lag spillovers from liquid macro and sector indicators into exposed small-cap stocks, but this advantage disappears at 100 bps when turnover and microstructure noise dominate. At 40 days, pure beta works for a different reason: slower macro repricing overtakes the firm-specific pure-alpha channel. The strongest conservative row is pure beta with GPT-4o mini sentiment, a Student-t target, a 40-day holding period, and risk parity allocation, reaching Sharpe 2.33 at 100 bps. The results suggest that stock-selection regime and allocator choice matter at least as much as the sentiment model, and that separating firm-specific and macro-exposure triggers is more informative than requiring both to fire simultaneously.

q-fin.PM↗

Heat transport in driven quantum systems: Comparison between the Floquet-Redfield equation and the master equation in the instantaneous eigenbasis

We provide a comprehensive study of heat transport in periodically driven quantum systems using a master equation approach in the Floquet basis. Starting from the exact, formal expressions, we obtain, within the weak coupling regime of system-bath coupling, a generalized Floquet-Redfield master equation which does not involve Markovian or secular approximations and the corresponding formula for the heat current. From this approach we derive the standard, full secular master equation. Moreover, we present a master equation for the driven spin-boson model in the instantaneous eigenbasis. The different approaches are compared by applying them to the steady-state heat transport in the driven spin-boson model. An analytical solution is provided for the master equation in the instantaneous eigenbasis which reproduces the numerics and explains the features of the heat current as a function of the drive parameters.

quant-ph↗

Universal magic state concentration

Magic plays a dual role in quantum computation: it promotes stabilizer dynamics from efficient classical simulability to computational universality, but also challenges fault-tolerant architectures, since non-stabilizer operations are harder to protect against noise. Magic state distillation addresses this issue, yet existing protocols remain largely tailored to specific assumptions about the input or noise model and, more fundamentally, no general information-theoretic theory of optimal magic-state conversion currently exists. Here we develop such a characterization for pure qubit states through universal magic state concentration: a fixed stabilizer protocol that converts a few copies of an unknown pure non-stabilizer qubit state into an exact target magic state. Motivated by the impossibility of exact $T$-state concentration, we build six- and eight-copy protocols producing an exact $\mathrm{CCZ}$ state, with six copies being minimal. Their success probabilities are governed by the linearized order-three stabilizer Rényi entropy $M^{\mathrm{lin}}_3$. We show that this connection is structural: for up to nine input copies, $M^{\mathrm{lin}}_3$ fully determines the success probability of every universal Clifford-invariant stabilizer protocol. Remarkably, this characterization persists asymptotically: our constructions achieve optimal rates among universal single-output protocols and remain optimal up to logarithmic factors among arbitrary stabilizer protocols. As a corollary, we show that any unknown pure non-stabilizer state suffices for universal quantum computation via probabilistic $\mathrm{CCZ}$-state injection. Together, these results identify the stabilizer Rényi entropy as a fundamental operational quantity in magic state distillation.

quant-ph↗

On the Acceleration of Pulsar Timing computations using Normalising Flows and Parallelisation

Single-Pulsar Noise Analysis (SPNA) and Gravitational Wave (GW) searches done on Pulsar Timing Array (PTA) datasets have everlastingly suffered from the computational bottleneck arising due to high dimensionality and multi-modality of the PTA likelihood landscape, along with strong correlations amongst various single-pulsar noises and ensemble-level common noise processes. We addressed this outstanding issue by employing a Normalising Flow-based Preconditioned Monte-Carlo sampling technique implemented in the POCOMC package, for the first time on PTA-specific computations, and comparing the achieved acceleration with the widely used PTMCMCSAMPLER and DYNESTY packages. We further investigated the acceleration achieved via parallelisation over an increasing array of communicating nodes on a high-performance computing (HPC) resource, by employing the PARALLEL_BILBY architecture with DYNESTY. We tested the acceleration on realistic long baseline simulated datasets with SPNA and Common Red Noise (CRN) analysis. We found PARALLEL_BILBY to be the most efficient in parallelisation, achieving a runtime of ~10min and ~100min with 16 nodes for spatially uncorrelated and Hellings and Downs-correlated CRN searches, respectively. POCOMC outperforms in single node performance requiring only ~10h for correlated search. PTMCMCSAMPLER was found to be the least efficient. We envisage POCOMC to be of great importance for PTA analyses, without requiring any GPU or HPC support, while also performing ensemble-level GW searches within a manageable time span. These results have everlasting implications with increasing data volumes and need to incorporate more complicated models, which were otherwise beyond reach due to the associated computational costs.

astro-ph.IM↗

MAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy

Vision-Language-Action models (VLAs) integrate visual perception, language instruction, and action generation into end-to-end policies across heterogeneous architectures. However, enabling VLAs to self-evaluate their action generation reliability without external supervision remains a major challenge. Existing methods either rely on expert annotations or estimate uncertainty only from output statistics, largely ignoring internal signals. In this work, we observe that internal visual modality entropy exhibits consistent distinctions between successful and failed tasks across heterogeneous VLAs. Although VLAs' architectures differ in their action generation, we show that they share a common latent action generation abstraction evolving under visual perception, language instruction, and State Input, which we formulate as a Conditional Generative Markov Chain. Based on this formulation, we propose MAE (Markov Attention Entropy), a self-evaluation framework that directly converts internal attention signals into architecture-aware reliability scores, and introduce LIBERO-Reflect, a 4,000-episode benchmark combining 2,000 standard episodes and 2,000 challenging episodes across four subsets. Extensive experiments across heterogeneous VLA architectures and diverse scenarios show that MAE consistently outperforms state-of-the-art baselines on AUPR, AUROC, and FPR@95.

cs.AI↗

Fault-Tolerant Quantum Computation with Adversarial Errors

We prove a fault-tolerance theorem for quantum computation against adversarial noise. For every quantum circuit on $\bar{N}$ logical qudits of depth $\bar{T}$, we construct a fault-tolerant circuit on $N=\text{poly}(\bar{N})$ physical qudits of depth $\bar{T}\cdot\bar{N}^{o(1)}$, which is robust against an adversary who may arbitrarily choose and corrupt an almost-linear number $N^{1-o(1)}$ of physical qudits at each time step. This robustness significantly improves upon prior fault-tolerance theorems, which assumed corruptions were either local and stochastic, or else only act on a polynomially vanishing fraction of qudits. Our fault-tolerance scheme addresses a key bottleneck towards constructing quantum PCPs via the circuit-to-Hamiltonian mapping of Anshu, Breuckmann, and Nguyen (STOC'24). More fundamentally, our result demonstrates that fault-tolerant quantum computation remains possible under noise models that are global, worst-case, and non-Markovian over the full duration of the computation, directly countering concerns that correlated noise could fundamentally undermine quantum fault tolerance. Our construction is based on a new family of subsystem product codes we develop, which have large dimension and distance along with low-weight parity-checks, and which support transversal non-Clifford gates. We show how to perform single-shot fault-tolerant error correction on these codes using a Floquet-like procedure based on the local testability of classical tensor codes. We then obtain a universal fault-tolerance scheme using repeated code switching in a hypercubic qudit architecture. Finally, we recursively compose our scheme with itself to reduce an initially exponential qudit dimension down to a constant.

quant-ph↗

Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory

Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VLA) and world-action models (WAMs) increasingly master individual skills, yet the chain still fails: errors compound beyond the policy's ability to correct, and one subtask silently constrains the next. A promising pathway freezes the VLA and puts an LLM coding agent in charge: it plans in language, moves in free space with analytic primitives, invokes the VLA only for contact-rich segments, and writes adaptation into language memory. Yet applied to long horizons, this recipe breaks twice. (1) Its competence comes from whole-task exploration at test time, whose cost is exponential in the number of stages: if one stage needs T episodes, a K-stage task needs on the order of T^K, and a failure does not reveal which stage caused it. (2) It has no representation of transitions: the VLA primitive carries an exit but no entry condition, and a subtask can succeed in a form its successor cannot use. We present BATON to address both failures. Against (1), BATON makes the subtask the unit of exploration: each subtask is explored in the cheap short-horizon regime and its solution stored in memory; a long-horizon trajectory is then composed from these solutions rather than discovered whole. Exploration cost becomes linear (KT), and each failure is attributed to one stage. Against (2), BATON equips exploration with a transition-aware memory. Within a subtask, a verifier agent governs the invocation transition: the VLA is invoked only after the wrist view confirms the scene is ready. Across subtasks, a handoff transition restores an entry state disturbed by the predecessor's residue, and a lookahead transition selects the strategy whose outcome the successor can inherit. On the RoboMemArena benchmark, BATON improves task success by 37.7% and cumulative success by 29.7% over the SoTA.

cs.RO↗

Wrong Operator or Blind Design? A Reference-Free Diagnostic for Physics-Informed Coefficient Learning

Physics-informed neural networks and hybrid models infer PDE coefficients from noisy data. When a trained network returns one, no standard check says whether to trust it. We show what those checks report when the operator is wrong: one sensor aggregating several diffusion sources. On one parabolic benchmark at $2\%$ noise, the in-domain error is $1.4$ times the noise while the identified diffusivity settles $30\%$ off. Every least-squares minimiser reaches that value, which drifts $27\%$ across windows; the network, whose objective is composite, settles $1.3\%$ away. The checks stay as silent when the design is blind to a rate of a richer operator, though the remedies are opposite. We develop a reference-free diagnostic, read in the physical parameter, not the weights, without retraining the network: an information-matrix test on the residuals, a heterogeneity statistic across window refits, and a Fisher-rank statistic on the design at the rates the single fit postulates. On the analytic head the specification test holds its pre-registered ceiling and rejects every misspecified replicate of both benchmark configurations, with a notch against a missing reaction term. The rank statistic is exactly zero only where the design is blind; a wrong operator confined to that mode leaves the specification test mute, and the rank statistic says so before any fit. The window reading exceeds its ceiling by one seed in thirty. A network frozen at its minimum returns the same verdicts; one stopped short rejects as a wrong operator would.

cs.LG↗

Frozen Memory Is Not Enough: Rethinking External Memory as Extraction

Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration with the backbone. Parametric adaptation is efficient at inference time, but entangles knowledge with model weights and can be hard to update, audit, or transfer. Engram-style hashed memory occupies a middle regime: it stores learned information in an external, addressable table, yet consumes that table through a small learned reader. This raises a basic question: when such a memory is moved across backbones, what matters more, the frozen memory itself or the target-side reader? We study this question through cross-model frozen-memory extraction, in which a memory trained on a source model is frozen and attached to a different target model, with only a lightweight reader trained. Ablations show that learned memory content and correct addressing both matter, but the transferred table becomes useful only through a reader aligned to the target model. In downstream question answering tasks, a dual-layer, four-branch reader nearly closes the gap between same-model and cross-model reuse, achieving an average score of 38.8 under our controlled evaluation protocol. Moreover, when the provider reader is directly compatible with the target interface, the frozen artifact can provide substantial utility without target-side training, while optional reader adaptation yields further improvement. These results suggest that Engram can serve as a reusable external knowledge artifact, provided that the target has access to a compatible reader interface; target-side adaptation can further improve alignment when direct reader reuse is insufficient.

cs.CL↗

Strong failures of club guessing at the successor of a regular cardinal

Starting from a model of $\mathrm{ZFC}$, we force the simultaneous failure of the Very Weak Club Guessing and the $\mho$ principles at $S_κ^{κ^+}$, for any regular cardinal $κ$. Our results are obtained by means of a forcing iteration technique, due to Krueger, that incorporates models as side conditions. At $ω_1$, the failure of these principles is a well-known consequence of $\mathrm{PFA}$.

math.LO↗

An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model should answer at that level rather than leak, evade, or refuse. We study this specification problem in a controlled LoRA-GRPO RWKU setting, comparing four reward designs that span lexical suppression, anti-refusal shaping, rubric-based broad answering, and an explicit refusal contrast, with and without SFT warm-up. The experiments show that optimization success is not equivalent to behavioral unlearning: RWKU forget scores, held-out completion audits, and training dynamics can point to different conclusions. We trace these disagreements to reward-hacking endpoints, policy-support limits in GRPO, benchmark probes that miss endpoint changes, and a rubric reward that selects broad-topic answering with low semantic leakage under held-out evaluation.

cs.LG↗

MoRAX: Mobility-based Representation Augmentation for Geospatial Foundation Models

Geospatial Foundation Models (GFMs) are emerging as a powerful paradigm for learning semantically rich and geographically consistent visual and physical representations. However, their reliance on Earth-observation (EO) data leaves information about human activity largely underrepresented. Human mobility data reveals the functional and relational structure between regions that is missing from EO data, but is often limited only to the city where it is observed, making it challenging to use for transferable urban representation learning. We introduce MoRAX, a lightweight framework for augmenting geospatial embeddings with functional structure derived from human mobility. MoRAX preserves the coverage and consistency of a GFM while providing information about the functional connectivity among urban regions, permitting zero-shot deployment in unseen cities with or without available mobility data. Across four target cities spanning two countries, the MoRAX teacher model, which observes mobility, consistently outperforms GFMs and strong urban representation baselines in eight socioeconomic and environmental prediction tasks. Meanwhile, the student model, which never takes mobility data as input, approaches the teacher in performance on most tasks. Transfer results across countries further demonstrate that modulation conditioned on mobility flows provides a general mechanism for grounding geospatial foundations in the human dimension of cities.

cs.LG↗

Asymptotic Entanglement Hiding under Stabilizer Restrictions

Entanglement is central to quantum information processing, while stabilizer operations underpin fault-tolerant quantum computation. We ask how much entanglement remains visible or distillable under stabilizer restrictions. We quantify stabilizer-visible entanglement by restricting the measured relative entropy of entanglement to stabilizer measurements, thereby obtaining converse bounds on entanglement distillation under stabilizer operations. We demonstrate magic-free asymptotic entanglement hiding: we construct explicit convex mixtures of pure stabilizer states on $N$ qutrits per party whose unrestricted visible entanglement and LOCC-distillable entanglement both grow as $Ω(N/\log N)$, while their stabilizer-visible and stabilizer-distillable entanglement vanish as $N\to\infty$. Thus, an unbounded amount of LOCC-distillable entanglement carried by stabilizer states can become asymptotically invisible and undistillable under stabilizer restrictions. We further prove that stabilizer-visible entanglement is $O(1)$ with high probability for Haar-random pure states despite extensive unrestricted visibility, and vanishes uniformly over entangled Werner states as the local dimension grows through odd primes. These results reveal a fundamental separation between entanglement and magic as resources, exposing intrinsic limits on entanglement extraction using stabilizer operations.

quant-ph↗

Harness Continual Learning: Continual Adaptation Beyond Model Parameters

Continual learning has largely been model-centric, treating model parameters as the state that changes with sequential experience. Modern agents can also adapt through a harness of prompts, memories, tools, skills, and routing rules. Because these contents jointly shape later execution, a harness update can disrupt previously reliable behavior even when the model is frozen. This raises a new question: how can an agent continually improve its state outside the model while retaining behavior acquired earlier? We formulate Harness Continual Learning (HCL), a new continual learning paradigm in which the harness evolves around a frozen foundation model, and define the resulting loss of earlier behavior as harness-level forgetting. We instantiate HCL with four execution-facing components: the Task Interface, Experience Memory, Capability Map, and Adaptive Router. We further introduce guarded harness evolution to separate update generation from state commitment. A Continual Optimizer proposes candidate harnesses from post-execution feedback, and a Continual Evaluator commits the resulting candidate harness only after checking current improvement, historical retention, and validity. Experiments on textual reasoning, multimodal perception, and open-world interaction demonstrate capability accumulation and failure recovery, with relative gains exceeding 10% over corresponding baselines in multiple settings. Component ablations assess the contribution of each harness component, while controlled retention sweeps reveal measurable harness-level forgetting and show that the stability--plasticity trade-off can be explicitly adjusted.

cs.LG↗

Twisted magnon frequency combs in ferromagnetic nanorings

We systematically investigate the emergence of twisted magnon frequency combs (tMFCs) and their higher-order modes arising from strong nonlinear coupling between vortex-core gyration and azimuthal spin-wave modes in ferromagnetic nanorings. The comb spacing is set by the gyrotropic frequency, which is controlled by both the size of the central hole and external magnetic fields. Remarkably, for the larger hole diameter (50 nm), an additional magnon mode emerges, leading to additional tMFC families. We also demonstrate that the selection rules still hold for different nanorings. In addition, the external in-plane magnetic field provides an effective means to tune the tMFC, while the response strongly depends on the nanostructure geometry. The nanodisk shows an approximately symmetric response under field reversal, whereas the response of nanorings depends strongly on the size of the central hole. For a small hole diameter (5 nm), the low-field response becomes asymmetric, and the tMFC spacing increases with field magnitude over the higher field branches. A larger hole diameter (50 nm) raises the gyrotropic frequency, yielding a sparser sideband structure near the drive frequency. Our results show that ferromagnetic nanorings support geometrically and magnetically tunable tMFCs.

cond-mat.mes-hall↗

One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows

Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a plausible response or valid tool call: agents must gather missing information over multiple turns, follow domain policies, coordinate dependent tools, and realize the correct persistent state transition without collateral effects. In this paper, we introduce Thinkingbox, a sandbox for tool-agent-user interaction that provides isolated MCP-compatible tool sessions, complete execution traces, and outcome evaluation over terminal backend state. Built on this sandbox, Thinkingbox-bench contains 507 policy-conditioned workflows across business scenarios, including retail, hospitality, auto insurance, neobank internal IT, and consulting IT/HR support. Each attempt is evaluated by task-specific executable checks that accept valid trajectories while rejecting wrong, missing, or extra effects; designated tasks additionally check required properties of the final response. Our experiments reveal that even the strongest proprietary and open-weight models show steep reliability drops: Claude Opus 5 falls from 66.50% pass@1 to 47.53% pass^20, and Kimi-K3 from 57.37% pass@1 to 17.60% pass^20. Moreover, many failed trials terminate cleanly after valid state-changing actions, so response- or tool-call-level signals poorly proxy end-to-end completion. Thinkingbox-bench reveals a large gap between occasionally finding a successful trajectory and reliably completing stateful business tasks.

cs.CL↗