Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 649 records · Page 36Linked to original sources

quARtet Marker: A 3D-Printable Multi-Tag Fiducial for Robust Near-Frontal Pose Estimation

Robotic manipulation of labware is difficult when transparent or reflective objects must be identified and localized. Coded planar fiducials are a practical retrofit: easy to print, they leave the marked face flat and graspable. Yet a single planar tag is least reliable in near-frontal views, where perspective cues fade. Non-planar geometries restore those cues but intrude on the flat face that a parallel-jaw gripper must contact. Our idea is to tilt multiple tags within one compact footprint, so that each tag is seen at a non-frontal angle even when the marker faces the camera. We propose the quARtet marker, a 3D-printable fiducial embodying this idea: all detected corners of its four tilted AprilTags enter one Perspective-n-Point solve, and a shared configuration defines the fabricated geometry and the detector model. Because tilting consumes flat area, its three layouts trade pose-estimation consistency against graspability. In robot-referenced, same-setup fixed-camera experiments, all three layouts reduced the mean frontal orientation error from 2.18 degree for a single planar tag to 0.24-0.47 degree and the root-mean-square position error from 1.50 to 0.17-0.20 mm. A robot-mounted-camera pose-hold test confirmed this separation under closed-loop visual feedback. In swing-down trials under identical conditions, the two layouts with flat contact strips retained the object, whereas the layout without flat strips slipped about a hundred times more than a single planar tag. For the tested conditions, the results support a rule: the layout without flat strips when pose-estimation consistency dominates, a layout with flat strips when the marked face must remain graspable.

cs.RO↗

Theory of Equivalent Tokamaks for Characterizing Turbulent Transport in Quasi-symmetric Stellarators

It is well known that quasi-symmetric (QS) stellarators are isomorphic to tokamaks in terms of their neoclassical-transport properties, and the corresponding transport coefficients can be calculated in the same manner as in tokamaks. However, less is known regarding the turbulent-transport properties of QS stellarators, e.g. the transport coefficients from the ion-temperature-gradient (ITG) mode. In this work, a systematic theory of the ``equivalent tokamaks'' for QS stellarators is presented based on the local gyrokinetic formulation and the near-axis expansion theory. It is shown that to zeroth order in the minor radius, the equivalent tokamaks can be chosen to have circular flux surfaces and can be characterized by three geometric quantities: the aspect ratio, the rotational transform, and the magnetic shear. To achieve first-order accuracy, however, not all QS stellarators have equivalent tokamaks, but good approximations can be found for some cases either as global or local equilibria. Local and global gyrokinetic simulations of ITG transport are performed for a selection of QS configurations, and quantitative agreement in the turbulent transport levels is found between the stellarators and their equivalent tokamaks.

physics.plasm-ph↗

Overlapping Subcritical Bubbles: Free Energy and Lifetime

Subcritical bubbles can form an appreciable population during weak first-order phase transitions, but are usually treated as isolated fluctuations. This raises the question of whether spatial overlap between neighboring subcritical bubbles can modify their evolution. We address this question by combining analytic free-energy calculations for composite Gaussian profiles with Langevin simulations of overlapping configurations. We find that the overlap lowers the free-energy cost and generally increases the lifetime of subcritical bubbles, with the enhanced persistence potentially feeding back on their abundance. Such overlap can therefore generate collective effects and should be incorporated into kinetic descriptions of subcritical-bubble populations.

hep-ph↗

Mentor-Initiated Asymmetric Bidirectional Quantum Teleportation Protocol for Arbitrary Qubit States

In this paper, a mentor-initiated bidirectional asymmetric quantum teleportation protocol is proposed. Unlike existing schemes that restrict the two-qubit state transmitted from Bob to Alice to a Bell-like form $α|00\rangle+β|11\rangle$, our protocol allows an arbitrary two-qubit state to be teleported from Bob to Alice within the mentor-initiated framework. The required entangled resources are constructed using standard quantum gates, and the protocol is illustrated through its quantum-circuit implementation in Qiskit. The proposed protocol is further simulated using Qiskit, and the simulation results show probability distributions very close to the theoretically expected values for the chosen input states.

quant-ph↗

RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations

An AI companion that talks with someone for months should come to understand them. It should know who they are, remember what they said, and recognize when something from the past matters now. Testing this needs real conversations, but real conversations are private, so existing benchmarks use invented people and invented questions. We release RealCompanion, ten real relationships between people and an AI companion, with 27,218 messages over up to 120 days. For each person, we release the full conversation, a profile, a persona, chat test items, and question test items. Every label points to the messages that support it, and every chat label comes with the reasoning that produced it. The real data shows three things. First, people rarely refer back. Only 3.4% of their messages depend on something said earlier, and when one does, the earlier message is usually far away (a median of 2,157 messages back). Averages hide this. Looking at the most recent messages finds the needed one 95.9% of the time overall, but only 2.2% of the time when it is far back. Second, AI systems cannot tell when the past matters. The detectors we tested barely beat chance on real messages, and when the same earlier messages are labeled "memories" instead of "earlier messages", models bring up the past 10 to 14 percentage points more often, even when nothing from the past is needed. Third, AI systems read more into a person than the person revealed. Three agent systems rebuild each persona equally well (F1 0.71). They see the person, and then imagine more. Understanding a person depends on knowing when their past matters and where what they shared ends, and only real conversations can test it.

cs.AI↗

Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2

Should low-precision transformer inference use stochastic rounding (SR) or round-to-nearest (RN)? The answer depends on where in the network you look. We isolate this effect by holding the numerical format fixed and varying only the rounding rule at individual operation sites. To enable experiments at freely chosen precisions, we extend the PRISM vectorized rounding library to arbitrary virtual precision via a variable-precision stochastic rounding (VPSR) algorithm, proving that the rounding decision is evaluated exactly in hardware floating point. We develop two analyses providing complementary insight into this site-level trade-off. First, a probabilistic forward-error bound for linear projections shows that SR's error envelope grows as $O(\sqrt{n} u)$ in reduction length $n$, versus $O(n u)$ for RN, a gap that widens rapidly at low precision and is most pronounced in the long multilayer perceptron (MLP) down-projection. Second, a second-order decomposition of expected cross-entropy loss change at the output softmax into signed drift, drift curvature, and a Fisher-weighted variance penalty reveals why the two sites behave oppositely: MLP noise is predominantly a uniform logit shift to which softmax is invariant, so SR's variance is largely discounted; head noise is non-uniform across the vocabulary and is not. On DistilGPT-2 at $t=6$ significand bits, observations match theory: SR in the MLP raises perplexity to 1.15x the full-precision reference, versus 2.21x for RN. At the language-model head, the ordering reverses because SR introduces non-uniform variance, whereas deterministic RN carries none. In a mixed-precision configuration (MLP output at $t=6$), assigning SR to the MLP and RN to the head brings perplexity within 1.10x of the full-precision reference, a 28% reduction over matched-bit RN.

cs.LG↗

UniWAM: Unified World-Action Model

Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semantic understanding and reasoning under distribution shifts. We introduce UniWAM, a unified architecture that integrates a physical reasoner, a world generator, and an action predictor to jointly learn semantic understanding of the physical world, visual generation, and action prediction. To ensure the quality of the training data, we developed a rigorous data cleaning and annotation pipeline for both human egocentric data and robot data. To adapt the vision-language component to embodied tasks while preserving its inherited language capabilities, we represent low-level actions in natural language and introduce a pre-training recipe that assigns complementary supervision from visual question answering (VQA) data, human egocentric data, and robot demonstrations to the appropriate model components. During post-training, future visual noise augmentation reduces reliance on precise future predictions, while history-conditioned flow matching uses encoded action history to initialize action generation. Together, these designs significantly reduce denoising steps while maintaining performance. UniWAM achieves state-of-the-art (SOTA) performance across multiple evaluations, including in-distribution performance, robustness, generalization, instruction following, and long-horizon task execution. Furthermore, we uncover a log-linear scaling law of unified human-robot co-training, demonstrating the effectiveness of large-scale pre-training on a mixture of human and robot data.

cs.RO↗

Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval

Memory-augmented large language models must decide which memories to retain, and recent systems do so by estimating each memory's effect on task performance. However, these estimates rely entirely on retrieved memories. When a memory is never retrieved, store-level interventions produce identical outcomes, leaving its utility unidentified. This is a retrieval-level positivity violation, invisible to diagnostics that examine only memory operations. We introduce Causal Memory Policy (CMP), a causal framework that restores identification by intervening on retrieval itself, reserving a fixed number of context slots for memories sampled with known propensities. CMP estimates memory utility by self-normalized inverse propensity weighting under a balanced assignment design. We prove the causal factorization of memory utility through retrieval, the unbiasedness and exact variance of the estimator, and the optimal decision rule under irreversible operations. Empirically, identification fails for 54% of required memories on LongMemEval and 67% on LoCoMo, and the failure persists in a deployed memory system. CMP improves discrimination between required and non-required memories from 0.54 to 0.66 AUC. Finally, we show that identified memory utility alone is insufficient for retention decisions: per-query utility reaches 0.78 AUC on the query for which it is estimated, yet no aggregation available to a retention policy predicts a memory's value on unseen queries. Code is available at: https://anonymous.4open.science/r/cmp-release-D0C3/.

cs.AI↗

Electron and monopole properties in a dark matter model

Monopoles have been a subject of much theoretical and experimental research since they were proposed to symmetrize Maxwell's equations. No experimental signature of their existence has been detected. A model generalizing QED with dark massless photons and dark monopoles was proposed not long ago. In this model, by introducing a mixing interaction between the two sectors, leptons and monopoles are transformed into into dyons. This particular behavior of the electrons determines a new path to study the physics of monopoles. Here we generalize the Electroweak theory in the same way by incorporating a dark sector with massive photons and monopoles, and analyze experimental consequences for electrons in both models. The mass of the dark photon plays an important role regarding observability.

hep-ph↗

Feasibility of Simultaneous Input-Output Constraints for Tracking in a Class of LTI Systems: Part I

This paper addresses simultaneous input and output constraint satisfaction for a class of multi-input linear time-invariant systems with state feedback and integral action. A control barrier function (CBF)-based governor modifies the reference command while preserving the nominal feedback controller. Necessary and sufficient conditions are derived under which the governor is feasible at every state of a prescribed operating set and ensures simultaneous input and output constraint satisfaction, forward invariance, and bounded closed-loop solutions. The paper also shows that these conditions need not be met for certain choices of the CBF governor parameters due to conflicts between input and output constraints. A systematic design procedure is proposed that first searches for feasible free parameters and, if necessary, proposes a relaxation in the input constraint that restores feasibility. A companion paper provides numerical examples and counterexamples illustrating these results.

eess.SY↗

Budgeted Cache Repair for Cross-Context KV-Cache Reuse

Cross-context KV-cache reuse predicts a shared segment's keys and values under a new prefix instead of recomputing them, and has been reported to do so without quality loss. We find otherwise, and identify two problems. (1) A hidden cost: on MMLU and GSM8K, reuse costs substantial accuracy. (2) A decision at the wrong unit: no rule for deciding whether to reuse a cache removes that cost. What does help is choosing which parts of the cache to recompute, and the value of choosing well falls as the unit of choice grows: informed selection removes 49.5% of the cache error beyond chance at single rows (one token's keys and values), 10.6% at 64-token chunks, and nothing at the level of whole calls. Budgeted Cache Repair (BCR) acts at the unit where selection still pays. It drafts two tokens from the assembled cache, ranks cache rows by the attention those tokens pay them, and recomputes a fixed budget of rows exactly, in one of three layouts. The cost is paid rather than predicted away, and the draft that fails as a gate succeeds as a selector. BCR restores GSM8K to dense-prefill accuracy while still serving most calls from cache, and its best layout outperforms every reuse baseline's mean in the reference grid. The draft also beats a coin-flip selector at the same budget - a control prior evaluations lack.

cs.AR↗

Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations

Large language models (LLMs) in production systems face prompt injections, trojans (backdoors), and manipulation of automatic quality metrics. This thesis develops models, methods, and algorithms for evaluating and improving LLM robustness to adversarial input sequence variations. We propose R_stab(f), a generative robustness metric based on the Jensen-Shannon divergence between per-step output distributions under small input perturbations. For localized attacks we prove V(h) <= 1 - R_class(h), where R_class(h) is the probability that a decision operator h keeps its decision under small perturbations. For non-localized attacks we propose a calibrated empirical model. For LLM-as-a-Judge systems we develop ASA, an adaptive evolutionary black-box attack that reaches an attack success rate (ASR) of up to 73.8%, with transfer between open models up to 62.6%. On Trojan Detection Challenge 2023 data (Pythia-1.4B), surrogate triggers reach REASR ~0.99 while recall of the true triggers is ~0.17 against a baseline of ~0.14. On SaTML CTF 2024 we systematize four classes of bypasses of multi-layer defenses, which reduce the ASR from 90% to 15-25%. Committees of 5-7 heterogeneous models reduce the ASR for Gemma-3-4B by 47-55 percentage points, to 19.3% with 7 models. For agentic systems based on the Model Context Protocol (MCP), we propose AttestMCP, which attests tool calls with HMAC-protected packets at under 0.1 ms per call, and the Commit Boundary isolation pattern. On the MCPBench benchmark of 847 scenarios they reduce the average ASR from 53.7% to 12.4%. The methods are implemented in the JudgeGuard and TrojanArmor software suites and the MCPSec module.

cs.CR↗

Single-photon addition to multimode quantum fields at telecom wavelengths using lithium niobate and its doped variants

Multimode spectro-temporal non-Gaussian quantum states of light play a crucial role in continuous-variable quantum information processing. In particular, the non-Gaussian quantum states generated through a single-photon addition (SPA) to multimode quantum fields at telecommunication wavelengths constitute a useful resource for quantum technologies. We present a detailed theoretical analysis and numerical simulations of single-photon addition (SPA) and mode-selective single-photon addition (MSSPA) to a multimode quantum field at telecommunication wavelengths. The analysis is performed for congruently grown lithium niobate (LN), $5\%$ MgO-doped lithium niobate (MgO:LN) and various concentrations of ZnO-doped lithium niobate (ZnO:LN) bulk crystals. We investigate MSSPA using type-II parametric down-conversion (PDC) and SPA using the type-I PDC processes. We further present spectro-temporal mode-selective SPA at the telecom wavelength using $5\%$ MgO:LN for temperatures ranging from $0^\circ C$ to $15^\circ C$. In addition, we perform simulations to investigate SPA based on the noncollinear type-I PDC configuration. These theoretical results identify optimal experimental conditions and phase-matching parameters to achieve SPA and mode-selective SPA for LN-based crystals, providing a potential platform for non-Gaussian quantum state engineering at telecommunication wavelengths, mainly at $1550$ nm.

quant-ph↗

Invariant Rings of Adjacent-Quadratic Triangular Derivations

Let $k$ be a field of characteristic zero and $D_N(a_n)=a_{n+1}a_{n+2}$ the adjacent-quadratic locally nilpotent derivation on $A_N=k[a_0,\dots,a_N]$. We determine the invariant ring at $N=5$, where deleting $a_0$ gives the polynomial base kernel $B^{D_5}=K=k[p,q,J_1,J_2]$, so localization gives the polynomial-line invariant algebra $(A_5^{D_5})_{pq}=K_{pq}[W]$ with $W=F_0/(p^2q^4)$, but this does not by itself determine the global kernel. An invariant $H$ of weight $24$ and $a_0$-degree $2$ lies outside $K[F_0]$, and adjoining it gives the six-generator hypersurface $C_5$; the obstruction is explained by algebraic residue behavior defeating naive denominator descent. An invariant $F_{46}^{int}$ of weight $46$ and $a_0$-degree $4$ lies outside $C_5$ and satisfies $p^2F_{46}^{int}\in C_5$. To pass from this candidate to the full invariant ring, a second slice chart arising from $f_{13}$, together with exact $q$- and $p$-saturation, yields $A_5^{D_5}=C_5[F_{46}^{int}]$, hence finite generation, with a codimension-two complete-intersection presentation and explicit Hilbert series. At $N=6$ the transported $B$-level ring is exact, while the $A_6$ analysis is established through total degree $13$ and the full $A_6$ invariant ring remains open.

math.AC↗

Imagine the Future, Internalize the Gist: Efficient VLA Reasoning via Internalized Spatiotemporal Imagination

Vision-language-action (VLA) models increasingly incorporate intermediate reasoning to improve robotic manipulation, yet existing approaches primarily reason about observed states without explicitly anticipating future scene evolution. Extending such reasoning to explicit future rollouts at every inference step, however, introduces substantial computational overhead. We propose IG-VLA, a VLA reasoning framework that enables models to imagine the future and internalize the gist. Our Latent Spatiotemporal Reasoning learns to imagine task-relevant future scene evolution directly in visual representation space, guiding action prediction without costly pixel-level video generation. To further reduce inference overhead, we introduce Scene Gist Memory, which internalizes reasoning-derived scene-behavior associations into a compact Scene Gist Token, preserving the benefits of future reasoning while bypassing explicit future imagination at inference. Extensive experiments on LIBERO, LIBERO-Plus, and VLABench demonstrate the effectiveness and efficiency of IG-VLA. On the LIBERO-Plus Language suite, both the reasoning and gist policies outperform the strongest baseline by nearly 6% in success rate. The gist policy also achieves up to 6.38x speedup over baselines, reducing inference latency from 1081ms to 169.5ms per action chunk on a single NVIDIA A6000 GPU. These results demonstrate that future spatiotemporal reasoning can be effectively internalized for efficient VLA deployment.

cs.CV↗

The zeroth stable homotopy groups of motivic spheres over the integers

The main result determines the zeroth integral Milnor-Witt stem of the motivic sphere spectrum in the Morel-Voevodsky motivic stable homotopy category of the integers. The component in weight zero is the Grothendieck-Witt ring of nondegenerate symmetric bilinear forms over the integers. Along the way, cellularity of connective Witt theory, as well as an absolute purity result for the η-inverted motivic sphere spectrum, is established over Dedekind domains of mixed characteristic.

math.AT↗

Multi-Task Evolution for Zero-Shot Cross-Problem Generalization using LLMs

Designing effective heuristics for diverse combinatorial optimization problems requires substantial expertise and repeated search. Large language models (LLMs) automate heuristic generation and refinement, but heuristic search typically depends on evaluation feedback from the problem being optimized. Generalizing to new problem definitions using only source-task feedback therefore remains a central challenge. We introduce MECo, an LLM-driven multi-task evolutionary framework for zero-shot cross-problem generalization. MECo maintains task-conditioned heuristic populations and uses a transfer gap based on cross-task population performance to guide their interactions. These interactions enable the transfer and recombination of heuristics. A complementary selection criterion then constructs a compact heuristic set by rewarding each member's additional coverage of source combinations. The selected set is applied to target problems without further search or adaptation. Experiments on 32 problem variants across vehicle routing (VRP) and flexible job-shop scheduling (FJSP) show that MECo achieves the lowest mean costs compared with eight automated heuristic design (AHD) baselines under the same budgets. On out-of-domain problems, it outperforms the strongest baseline in each family. Moreover, integrating the framework of MECo with different AHD methods improves their ID and OOD performance in both families, supporting its effectiveness across different methods.

cs.AI↗

Beyond Entropy: Self-Diagnostic Multi-Role Token Optimization for Video Reasoning

Reinforcement learning with verifiable rewards has substantially advanced multimodal reasoning, yet it remains fundamentally limited by ambiguous token-level credit assignment. While high-entropy token heuristics encourage possibility exploration, naively extending them to video reasoning tends to induce lengthy reasoning, as the model becomes overly reliant on high-entropy visual activations. Alternative approaches that rely on counterfactual-based visual token localization for credit assignment also tend to over-prioritize visual exploration at the expense of decisive reasoning cues for answer derivation, thereby exacerbating the interference from spurious visual nuances. Moreover, these methods employ static counterfactual strategies that fail to co-evolve with the policy during training. In this paper, we introduce DyCPO, a co-evolutionary framework that jointly optimizes reliable token selection and adaptive counterfactual intervention. It constructs a multi-role dependence metric to balance visual exploration and answer-relevance mining in token-wise contrastive learning, while suppressing exploration-only filler tokens and spurious visual noise. Rather than relying on static counterfactual priors, DyCPO dynamically derives counterfactual signals from the model's own successful and failed rollouts, enabling self-diagnostic analysis and co-evolution of the optimization objective with the policy. Extensive experiments on complex video reasoning and general video understanding benchmarks demonstrate consistent performance improvements, establishing DyCPO as a robust token-level credit assignment paradigm for multimodal reinforcement learning.

cs.CV↗