Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 397 records · Page 22Linked to original sources

Dyna-DINO: Efficient ViT Distillation Via Adaptive Representation Anchoring

Vision Foundation Models (VFMs) with Vision Transformer (ViT) backbones, such as DINOv2, have become essential for downstream tasks like object recognition and semantic segmentation. The immense computational requirements of backbones often necessitate distillation into smaller architectures for edge deployment. Feature-based knowledge distillation (KD) often suffers from the teacher-student gap; the student struggles to imitate teacher's complex feature map due to its limited capacity. To mitigate this bottleneck, we propose Dyna-DINO: Efficient ViT Distillation Via Adaptive Representation Anchoring, a training curriculum for ViT feature-based knowledge distillation. By utilizing the teacher's intermediate feature maps as a sequence of progressively more difficult targets, our curriculum allows the student to build a foundational representation before tackling higher-level abstractions. Our results demonstrate that this paradigm significantly accelerates convergence through adaptive difficulty selection across various student model sizes and dataset scales. With our curriculum, the Dyna-DINO distilled ViT-S achieves 90.1% accuracy on ImageNet-100, a +12.24% improvement compared with baseline. On ImageNet-1K, Dyna-DINO achieves +3.9% and +6.09% improvement for the instance retrieval task on the Oxford and Paris datasets, +1.93% improvements on semantic segmentation task, as well as meaningful performance gain on classification task. Furthermore, the curriculum enables 25.1% savings in training FLOPs and 21% savings in training time on ImageNet-100 by implementing early-stopping for teacher inference during the initial stages of training. Code is available at https://github.com/KevinZ0217/Dyna-DINO

cs.CV↗

A Law of Iterated Expectation Primer for Causal Inference

The g-formula is a foundational tool for identifying causal effects in observational data. This tool is based on the law of iterated expectation, a key mathematical identity in statistics. However, the notation with which the law of iterated expectation and the g-formula is expressed can be opaque to those with little background in statistics. We provide a primer introducing the law of iterated expectation, the integration notation used to express it, and its role for causal effect identification via the g-formula. Combined with the assumptions of causal consistency, positivity, and conditional exchangeability, the law of iterated expectation yields a causal standardization formula (the g-formula) in two nonparametrically equivalent forms: a non-iterative conditional expectation (NICE) form involving a single weighted average of conditional outcome means, and an iterative conditional expectation (ICE) form involving nested expectations. We illustrate both forms using three progressively complex numerical examples: a time-fixed example with a single binary confounder, a time-fixed example with discrete and continuous confounders, and a time-varying example with two timepoints. We provide clarity on what the law of iterated expectation is, how it is related to the g-formula, and how to gain intuition of its mathematical formulations in actual data examples that can be generalized to a range of settings.

stat.OT↗

Compressing History into Memory: Distilling Transformers into Recurrent Transformers

Transformers are AI's workhorse but their computational cost becomes prohibitive when processing long sequences. We target long-horizon streaming vision and robotics applications, where it is particularly impractical to store and maintain a history of observations. Recurrent Transformers address this limitation by maintaining fixed-size memory but their performance lags behind that of transformers operating over the full observation history. We argue that this gap does not stem from architectural limitations, but from differences in how these models learn to compress past information. Without access to an observation history, recurrent models must explicitly decide what to retain in memory at each step, a significantly harder learning problem. In this work, we propose a distillation approach that transfers the compression strategy of a classical full-history transformer to a recurrent variant. We enable this by designing a teacher model that explicitly compresses its observation history into a fixed-size bottleneck representation and directly supervise the student's memory with this bottleneck representation, effectively aligning the two compression mechanisms. We show that this approach allows to train a recurrent latent robotic memory with linear-time complexity on the Mem-RPE task while substantially narrowing the performance gap to full-history transformers. We additionally validate the same principle on streaming visual question answering (VQA) and observe improved recurrent predictions thanks to memory distillation

cs.CV↗

On $k$-limited domination: complexity and Cartesian products

A dominating set is called $k$-limited if every vertex in the set has at most $k$ neighbors outside it. The minimum cardinality among all $k$-limited dominating sets of a graph $G$ is the $k$-limited domination number, denoted by $γ_k^{\mathrm L}(G)$. We prove that, for every fixed integer $k\ge 2$, deciding whether a graph admits a $k$-limited dominating set of size at most $\ell$ is $\mathsf{NP}$-complete. In addition, a systematic study of $k$-limited domination in Cartesian products is initiated. In particular, we establish general lower and upper bounds for $γ_k^{\mathrm L}(G\square H)$, show that both are sharp, and derive exact values for several natural families of graph products.

math.CO↗

Hierarchical Quantum Logical Processor with Amortized Long-Range Connectivity

High qubit overhead is a key bottleneck for fault-tolerant quantum computation. Quantum low-density parity-check (qLDPC) codes offer high encoding efficiency but typically require non-local connectivity in every syndrome extraction cycle, incurring additional physical errors and implementation complexity. We introduce a hierarchical logical processor (HLP) architecture that implements a high-rate quantum CSS code with distance-$d_{0}$ rotated surface codes (RSC), allowing long-range connectivity to be used only once every $Θ(d_{0})$ rounds of surface-code syndrome extraction. HLPs introduce elongated RSC patches called shuttle buses. Using transversal hybrid-unit CNOT gates, a single shuttle bus can simultaneously couple to multiple standard RSC patches. This capability enables efficient level-1 syndrome extraction with suppressed level-1 error correlations and supports highly parallel logical Pauli measurements. We perform circuit-level simulations of several concrete HLP constructions and benchmark both logical memory and logical Pauli measurement performance. At a physical error rate of $10^{-3}$, a hierarchical logical processor can achieve 3-4 times higher qubit efficiency than the rotated surface code and 20-30 times shorter logical error-correction cycle times than the yoked surface code.

quant-ph↗

Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Representations

Refusal on a safety benchmark does not reveal how stable that behavior will remain after model updates. Benign downstream fine-tuning can weaken refusal, yet behavioral evaluations typically expose this fragility only after an intervention. We introduce SKIN-DEEP, a geometric diagnostic that examines the unmodified model's residual-stream activations. It compares aligned and base checkpoints to identify safety-separating directions, tests their behavioral relevance through ablation, and summarizes the layer-wise pattern in the Geometric Fragility Score (GFS). Across twenty-one instruction-tuned models, harmful requests and benign instructions exhibit a recurring low-rank separation pattern. Selected direction ablations weaken refusal, with the effective direction varying across models. In benign low-rank fine-tuning experiments, the initially safe model with the lowest score before fine-tuning has the lowest harmful-compliance rate when trained on the largest tested set of harmless examples. These findings connect representation geometry to subsequent behavioral susceptibility and support activation-based diagnostics as a complement to refusal tests. Our code is available at https://github.com/js-lee-AI/skin-deep.

cs.AI↗

Diffuse Supernova Neutrinos with Secret Neutrino Interactions

The Diffuse Supernova Neutrino Background (DSNB), an isotropic flux arising from the cumulative neutrino emission of all stellar core-collapse events throughout cosmic history, is expected to be detected by next-generation neutrino observatories. As DSNB neutrinos propagate over cosmological distances through the cosmic neutrino background (C$ν$B), they may undergo non-standard neutrino self-interactions ($ν$SI), leaving distinct spectral imprints on the observed flux. In this work, we investigate the impact of scalar ($ϕ$)-mediated $ν$SI on the DSNB within a full three-flavor framework that retains the complete PMNS structure. We consider four representative flavor-diagonal coupling structures--universal, $e$-, $μ$-, and $τ$-specific. The resonant scattering $ν_iν_k\toϕ\toν_jν_l$ off the lightest, relativistic C$ν$B state produces broad spectral depletion whose pattern depends on the coupling structure and the neutrino mass ordering, generating distinctive signatures across the six flavor fluxes. We compute the resulting event spectra at JUNO, Hyper-Kamiokande with gadolinium loading, and DUNE, and derive projected $3σ$ sensitivities in the $(m_ϕ,~g)$ parameter plane. We find that these experiments can probe couplings as low as $g\sim10^{-8}$ for $m_ϕ\sim100$--$300$ eV, surpassing existing bounds by up to a few orders of magnitude in the sub-100 eV mass range. Moreover, unlike the flavor-blind cosmological and supernova bounds, the DSNB sensitivity is flavor-discriminating, offering a unique opportunity to identify the underlying flavor structure of $ν$SI in the event of a detection.

hep-ph↗

Source-Free Detection and Impact Analysis of Compiler Optimization Problems in Mobile Applications

Mobile apps frequently suffer from frame drops, overheating, and excessive power consumption. While developers optimize algorithms and debug code, a critical bottleneck often goes unnoticed: native libraries compiled with low optimization levels (O0/O1 instead of O2/O3). Because these libraries execute without functional errors, the resulting performance degradation remains hidden in production apps. We present \textsc{OptDetect}, a source-free framework that detects compiler optimization problems directly from app binaries. \textsc{OptDetect} handles mixed optimization levels through binary disassembly, chunk-level classification, and weighted score aggregation, achieving 93.0\% accuracy on controlled datasets and 81.9\% on real-world datasets. Applying \textsc{OptDetect} to 21,972 native libraries from 830 top-ranked Google Play apps, we find that 30.5\% of libraries use low optimization levels, affecting 91.7\% of apps. Through case studies on 12 production apps, fixing detected issues reduces CPU instructions by 10-63\% (median: 20.5\%) for commercial apps and 15-58\% (median: 32\%) for open-source apps. Performance complaints decrease in 5 of 6 commercial apps, and ratings increase in 5 of 6. Further investigation reveals that widely-used third-party libraries are themselves distributed at low optimization levels, with 49.7\% of 1,073 libraries in a major repository exhibiting this problem. These findings show that compiler optimization problems are common, source-free detectable, and practically consequential in mobile app ecosystems.

cs.SE↗

Bottom quark electroweak dipole moments at a high-energy $μ-$collider

We study the sensitivity of a high-energy $μ-$collider with center of mass energy in the multi--TeV range in testing electroweak dipole interactions of the $b-$quark. We parametrize the relevant deformations in the language of the Standard Model effective field theory, where the dominant modifications arise at the $d=6$ level. We analyse $μ^+ μ^- \to b \bar b$ and $μ^+ μ^- \to b \bar b h$ scatterings, performing a study at the level of a fast detector simulation. Owing the chiral structure of the dipole interaction, the study of the $μ^+ μ^- \to b \bar b h$ process allows to enforce the stronger bounds on the Wilson coefficients of the $d=6$ operators. The limits that can be obtained surpass present and future bounds from EW precision measurements also improving upon the ones arising from the measurement of the $ΔF=1$ transitions $B\to X_sγ$.

hep-ph↗

MemAudit: Auditing Long-Term Agent Memory via Hidden User-State Recovery

Long-term memory promises LLM agents that grow more capable across sessions, maintaining an accurate, evolving understanding of the user that interaction forms. In practice, however, this memory is evaluated mostly through downstream behavior, such as later answers, personalization quality, or task success, which tests that understanding only indirectly and leaves the memory artifact itself largely unaudited. We argue that long-term memory should instead be evaluated as an auditable post-interaction artifact: after ordinary assistance, what structured user state can be reconstructed from the memory the agent leaves behind? We instantiate this view in MEMPROBE, a benchmark in which a memory-equipped agent assists simulated users, each carrying a hidden, taxonomy-anchored user-state bank, across a trajectory of leak-controlled tasks, after which that bank is reconstructed from the agent's resulting memory under both full-store and top-k access. Built on synthetic ground truth for efficient, scalable measurement, MEMPROBE spans 50 simulated users with 31 hidden dimensions each (1,550 recovery targets) and tests 5 representative memory systems. Testing state-of-the-art memory agents, we find that successful assistance and recoverable memory behave as distinct capabilities. Task completion nearly saturates, even for a memoryless baseline, while category-balanced recovery stays moderate (about 0.6) and drops further under top-k retrieval. MEMPROBE is the first benchmark to study memory recovery directly, reconstructing the user state a system retains and scoring it against ground truth. We see recovery as a concrete objective for future memory agents to optimize, and MEMPROBE as a step toward an environment where agents are trained to remember their users, growing more faithful the longer they know them.

cs.CL↗

Arbitrarily Loss-Tolerant Quantum Position Verification in a Single Execution

Quantum position verification (QPV) aims to certify the location of an untrusted prover, but faces two major obstacles: fundamentally, entanglement-based attacks and, experimentally, photon loss. A commitment-based modification introduced in Phys. Rev. Lett. 135, 260801 addresses both in sequentially repeated protocols. Its security analysis, however, relies on the sequential structure and does not extend to parallel repetition. We use a different proof approach to establish security of the commitment modification for a parallel BB84-based QPV protocol. Against bounded-entanglement adversaries, the acceptance probability decays exponentially in the minimum number $k$ of successfully committed qubits. The protocol tolerates arbitrary transmission loss and noise rates up to 3.7%, while allowing arbitrarily slow quantum communication. This yields a fully loss-tolerant, single-execution QPV protocol secure against bounded-entanglement attacks, removing transmission loss as a fundamental limitation on verification distance. We also revisit sequential repetition, correcting the treatment of conditioning on commitment and replacing the earlier conditional guarantee with an unconditional security bound, without the finite-size correction to the correctness gap. The resulting bound provides improved quantitative security parameters for experimental implementations.

quant-ph↗

The renormalization of the shell-model neutrinoless double-beta decay operator starting from effective field theory at leading order

In this work, we approach for the first time the task to perform a shell-model calculation of the matrix element for the neutrinoless double-beta decay, within a fully-consistent framework where the expressions of the nuclear Hamiltonian and of the decay operators have been derived through chiral perturbation theory. More precisely, the effective shell-model Hamiltonian and all transition operators have been constructed by way of the many-body perturbation theory, and then employed to calculate both spectroscopic properties of the nuclei involved in the decays under our consideration - namely 48Ca, 76Ge, and 82Se -, as well as the nuclear matrix elements of the electromagnetic and neutrinoless double-beta decays. We also present a study of the convergence properties of the calculated matrix elements in order to provide the elements for an estimate of the theoretical uncertainty.

nucl-th↗

TESLA-for-5G: Broadcast Authentication for 5G Networks Using TESLA

5G base stations broadcast unauthenticated system information (SI) that every user equipment (UE) reads during cell selection. This enables attackers to broadcast forged SI from a fake base station (FBS), deceiving UEs into camping on it. Prior approaches to address this issue usually require UEs to authenticate System Information Block 1 (SIB1) using digital signatures. This necessitates computationally expensive verification for every SIB1 reception, imposing a significant burden on resource-constrained UEs. We propose TESLA-for-5G (TF5), a broadcast authentication protocol for 5G SIB1 that combines TESLA with GG09 Schnorr-like identity-based signatures (IBS). In the steady state, TF5 enables UEs to authenticate each SIB1 message using a symmetric MAC and delayed key disclosure, eliminating the need for per-message digital signatures. Initial trust is bootstrapped during cell entry using a lightweight GG09 IBS over the TESLA parameters, avoiding certificate distribution overhead. We formally verify the security of TF5 in Tamarin under a Dolev -Yao adversary and demonstrate its favorable computation, communication, and storage costs through both an implementation on the OpenAirInterface 5G stack and trace-driven analysis.

cs.CR↗

Direct Action-Head Injection of A Grounded 3D Point Unlocks Spatial and Task Generalization

Vision-Language-Action (VLA) models leverage large-scale vision-language pretraining for flexible robot manipulation, yet at test time they remain brittle to changed object positions and to familiar scenes paired with different instructions. A growing family of methods addresses this brittleness by supplying the policy with grounding signals, such as 2D pixel coordinates for object localization and placement. However, we find that how the grounding signal is represented and injected matters more than the signal itself. In this work, we propose a lightweight module that represents the grounding signal in 3D and injects the resulting embedding directly into the action head. The module is a two-layer MLP and requires no changes to the VLA backbone or pretraining pipeline, yet it yields substantially larger gains than language- or visual-prompting alternatives. On LIBERO-PRO, our method improves the average success rate of GR00T-N1.6 from $31.2$ to $77.5$ under task perturbation and from $28.1$ to $60.2$ under position perturbation. Comparable gains are also achieved for $π_{0.5}$, demonstrating that the mechanism is backbone-agnostic across VLAs with diffusion-based action heads. We further validate the practical applicability with real-world experiments. Together, these results support our central finding: lifting adequate 2D grounding into 3D and injecting it into the action head enables spatial and instance-level task generalization in VLAs.

cs.RO↗

On the generic structures of the protocols for quantum auction and quantum summation and their relation

Secure multi-party computation (SMC) addresses the problem of jointly computing global functions of private inputs while revealing minimal information about individual data. Two prominent examples of SMC tasks are sealed-bid auction and secure multi-party summation. Existing schemes for quantum auction and quantum summation have largely been developed independently, motivated by distinct applications and employing different computational primitives. In this work, mutual reductions between the two tasks are established for single-item sealed-bid auctions with first-price payment. In one direction, an auction used as a black box computes a secure sum: each party bids a private random key drawn from its input, so the clearing price depends on the inputs only through their sum $S$. The reduction is perfectly private against colluding semi-honest parties and uses $O(S^2\log(1/δ))$ auction rounds, which is optimal, or $O(S\log(1/δ))$ under coherent access. It can be instantiated with an existing auctioneer-free quantum auction after a one-line change to its disjunction subroutine, which as published leaks occupancy counts. In the other direction, summation with randomized inputs yields threshold disjunctions, and a single search over composite keys locates the highest bid and a winner, revealing nothing else. Because each query consumes the bidders' encodings, a bidder can adapt its effective bid undetected, while a fixed query set restores sealed bidding. Combinatorial auctions and general payment rules lie outside this framework. On a superconducting processor, one instance of each direction is executed, recovering the winner for four-bidder instances with ties and the sum of private inputs from repeated auctions.

quant-ph↗

Pepti-drift: Scalable Safe-Active Peptide Generation Without Inference-Time Guidance

Therapeutic peptides are a promising drug modality, but their generation must satisfy multiple therapeutic constraints. We introduce BindSafe-PepBench, a fixed-budget benchmark that jointly evaluates target binding and four major safety metrics on the same generated candidates. We reveal that peptide length is a major confounder of joint binding-safety evaluation: longer peptides tend toward stronger predicted binding but less favorable predicted safety. This creates an apparent trade-off and can bias comparisons among models with different output-length distributions. We report absolute Safe-Active yield and exact-length-matched gains to distinguish generative improvements from output-length effects. High Safe-Active yield remains challenging, while the strongest multi-property methods rely on costly inference-time guidance. We therefore introduce Pepti-drift, a one-step generation framework that incorporates attraction toward target-specific binders and repulsion from liability-associated regions, requiring a single latent refinement followed by parallel decoding without inference-time guidance. Across 88 held-out targets, Pepti-drift achieves an 18.37% predicted Safe-Active yield while retaining positive exact-length-matched gains. The resulting gains are competitive with multi-property-guided baselines while requiring 468 times lower generation cost, enabling scalable and fair high-throughput peptide design.

cs.LG↗

Disentangling Continuous-Time Latent Dynamics: Identifiability of Latent SDEs via Diffusion Shifts

Causal representation learning for time series has developed strong identifiability results in discrete-time latent causal models, but identifiability in continuous-time latent stochastic differential equation (SDE) models remains largely open. We address this gap using environment-induced shifts in diffusion covariance. We study additive-noise latent SDEs observed through an unknown nonlinear diffeomorphism, with shared drift but environment-specific diffusion covariance. We show that two diagonal diffusion regimes with pairwise distinct coordinate-wise variance ratios identify the latent coordinates up to permutation, coordinate-wise scaling, and a possible constant shift, without any sparsity assumption on the drift. We first prove this result for linear Ornstein-Uhlenbeck systems and then extend it to general additive-noise latent SDEs. Under mild smoothness, the instantaneous drift-Jacobian causal graph is identifiable up to the same permutation. We propose a two-stage estimator for latent disentanglement and optional graph recovery; experiments on synthetic systems confirm the predicted identifiability boundary, and an application to Hardanger Bridge monitoring data illustrates the approach on real sensor trajectories.

cs.LG↗