Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Cross-species representation learning aligns mouse and human neural dynamics and tracks clinical drug efficacy

Preclinical models poorly predict human drug efficacy, particularly in neurological disorders. Neural activity offers a uniquely rich source of translational information because it captures high-dimensional variation in nervous-system function that can be measured in both animals and humans. However, its high dimensionality makes it difficult to distinguish conserved disease-related features from variation arising from species, recording modality and experimental context. Here, we test whether shared neural dynamics can be identified directly from electrophysiology data by learning representations organized by biological state rather than species. We develop a dual-rule contrastive learning framework that aligns corresponding mouse and human states while preserving separation between distinct phenotypes. This framework recovered conserved sensory-response structure across species and, in epilepsy, resolved distinct relationships between three mouse models and heterogeneous human patient populations. When treated animals were projected into a frozen cross-species representation, drug-induced movement towards the human-aligned healthy state retrospectively tracked known clinical efficacy across ten model-drug combinations including a disease-specific detrimental effect. The framework also identified shared disease-associated neural dynamics between Fmr1-knockout mice and human 16p11.2 copy-number variant carriers despite differences in genetic aetiology and recording modality. Together, these findings show the potential of cross-species neural representation learning to map heterogeneous human disease onto experimentally tractable preclinical states and assess whether interventions restore human-relevant circuit function.

q-bio.QM↗

SafeInferCom: Safe Inference-Time Compute via Verifier-Guided Mid-Generation Intervention for Robotic Task Planning

Large Reasoning Language Models (LRLMs) enable multi-step reasoning for robotic task planning, but continued reasoning can overwrite valid intermediate plans or leave constraint violations unresolved, reducing planning reliability and wasting inference-time computation. We develop an inference-time monitor that exposes and verifies intermediate plans without disrupting the original decoding trajectory. Building on this monitor, we propose SafeInferCom, a formal verifier-guided framework that preserves valid intermediate plans and directs error correction during generation. Experiments across multiple LRLMs and planning domains reveal reasoning-response inconsistency and limited self-correction under one-shot inference. SafeInferCom improves planning success and accelerates error correction relative to one-shot inference. When combined with iterative refinement, it further improves success while reducing token usage compared with refinement alone. We additionally evaluate SafeInferCom in VirtualHome and provide a real-world robotic-arm demonstration.

cs.RO↗

Online Adaptive Computation Reuse in Collaborative Edge Computing: A Two-Timescale Approach

Collaborative Edge Computing (CEC) is an efficient computing paradigm that enables neighboring edge servers to share computational resources with each other. Although CEC can enhance resource utilization, it still suffers from duplicate computations because nearby end-users often offload tasks with similar inputs. To improve system efficiency, the computation results of previously executed tasks can be cached and reused by subsequent tasks. However, time-varying task popularity and arrival rates require caching and scheduling decisions to adapt to demand changes, while frequent cache updates incur additional costs. To address this issue, this paper develops a two-timescale computation reuse algorithm for CEC networks. We formulate an optimization problem that jointly considers weighted response time and cache update cost, with result caching decisions determined at the frame level and workload scheduling, cache searching, and computational resource allocation adjusted at the slot level. Using recent workload observations, we construct a frame-level surrogate problem and decompose it into a caching subproblem and a scheduling subproblem. For the caching subproblem, we introduce marginal storage efficiency and incorporate cache update costs into a bisectionbased algorithm. Under bounded normalized marginal sensitivity, the algorithm achieves a near-optimal objective value for the single-BS caching subproblem when individual result sizes are small relative to the cache capacity and the relaxed solution is sufficiently accurate. For the scheduling subproblem, we utilize projected gradient descent and backtracking with warm starts across consecutive slots. Numerical results demonstrate consistent performance gains over benchmark schemes across diverse network and workload settings.

cs.NI↗

High-Rate Concatenated Quantum Error-Correcting Codes for Qudits

Qudits provide a larger local Hilbert space than qubits and have recently attracted interest as building blocks for efficient fault-tolerant quantum computation (FTQC). In this work, we extend many-hypercube (MHC) codes, high-rate concatenated quantum codes, from qubits to qudits of prime dimension $q$. We construct the level-2 and level-3 MHC codes with parameters $\left[\!\left[6^2,4^2,2^2\right]\!\right]_q$ and $\left[\!\left[6^3,4^3,2^3\right]\!\right]_q$, respectively, and generalize the level-by-level minimum-distance (LLMD) decoder originally proposed for the qubit MHC codes. In a qudit $X$-error model, we evaluate the code-capacity performance of the qudit MHC codes for dimensions up to $q=13$. We find that the error threshold increases monotonically with the local dimension $q$, reaching $10.0\%$ for $q=13$, nearly double the value of $5.1\%$ for qubits. Furthermore, for sufficiently large local dimensions, we observe a pronounced waterfall regime in which the logical error rate decreases substantially faster than expected from conventional distance-based estimates. To understand the origin of the observed qudit advantage, we analyze the decoding performance in terms of the entropy of the underlying $q$-ary noise channel. This analysis reveals a competition between the increased syndrome information available at larger $q$ and the simultaneous growth of the typical physical error weight at fixed entropy. These results demonstrate that higher-dimensional quantum systems can substantially improve the performance of high-rate concatenated quantum codes and motivate further development of qudit-based FTQC architectures.

quant-ph↗

When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM Quantization

Weight-only post-training quantization (PTQ) relies heavily on reconstruction loss minimization to preserve model quality at low precision. We show that the weights favored by minimizing this loss need not yield better model performance on new tasks. In fact, we find that lower reconstruction loss can even degrade model performance on the same calibration data. Our analysis further shows that weights with lower reconstruction loss on calibration data can have higher loss than other weights when the distribution of input activations changes. Motivated by these observations and our analysis, we propose Distributionally Robust Quantization (DRQ), a post-hoc refinement process that minimizes worst-case reconstruction loss over a constrained set of input activation distributions. DRQ refines the integer codes representing quantized weights within the existing quantization grid, keeping quantization parameters and inference operators unchanged. Extensive experiments show that DRQ improves models quantized by six representative PTQ methods, including AWQ, GPTQ, and ParoQuant, and delivers gains across both dense and mixture-of-experts large language models. These results establish DRQ as a general post-hoc refinement framework for weight-only PTQ, achieving better downstream performance without adding inference overhead.

cs.AI↗

Fourier Restriction Beyond Dimension Data: A Multiscale Moment Criterion

We establish an exact Fourier restriction criterion for a class of weighted Moran measures. Assuming uniformly bounded one-level frame bounds, we prove that, for every \(s>2\), the measure itself already detects the full extension estimate: \(L^2(μ)\to L^s(\mathbb R)\) extension holds if and only if \(\widehatμ\in L^s(\mathbb R)\), equivalently, if and only if the one-level Fourier moments \(G_{s,n}\) are summable, where $G_{s,n}$ is the $s$th moment of the nonzero Fourier coefficients of the $n$th digit law. The sufficient direction is obtained from a finite-frame stability estimate and its multiscale iteration, while necessity follows from disjoint resonant frequency grids. This gives, in particular, an exact formula for the critical restriction exponent and a precise criterion for endpoint attainment. The criterion reveals that the optimal restriction range is not determined by standard Fourier or dimensional data. We construct singular continuous spectral Moran measures with the same scales, digit cardinalities, exponential spectrum, and maximal nonzero Fourier coefficients. They also satisfy the same global Fourier decay bound, which is sharp along the same sequence of resonant frequencies, yet their critical restriction exponents are different. More generally, after fixing the summability rate of the maximal Fourier coefficients, we determine the optimal interval of possible restriction exponents and show that, at every interior critical exponent, the same fixed data allow both attainment and failure of the critical endpoint. These measures may simultaneously have $ \dim_{\mathrm F}^θμ=θ$ and $ D_q(μ)=1,$for all admissible $θ$ and $q$, and pointwise local dimension one everywhere on their supports.

math.CA↗

Multimodal Graph Retrieval-Augmented Sequential Recommendation via Collaborative Filtering Paths

Multimodal Large Language Models (MLLMs) have demonstrated strong potential for sequential recommendation through their ability to reason over complex multimodal data. However, existing approaches either rely solely on the target user's own interaction history, neglecting collaborative signals from neighboring users, or incur substantial computational overhead through repeated MLLM inference over long interaction histories. To address these challenges, we propose MGRASRec, a multimodal graph retrieval-augmented framework for sequential recommendation. MGRASRec injects collaborative filtering signals conditioned on the candidate item directly into the MLLM prompt by retrieving structured paths from a user-item interaction graph, extended via multimodal similarity to increase coverage beyond exact co-interaction overlap. This retrieval also surfaces the history items most relevant to the candidate at no additional cost, removing the need for recurrent summarization and keeping inference to a single forward pass per candidate. All components are unified into an augmented prompt for parameter-efficient fine-tuning of an MLLM. Extensive evaluations across three publicly available datasets validate the effectiveness of MGRASRec, achieving the best performance on all metrics with particularly strong gains in ranking quality.

cs.LG↗

SteerCast: Retrieval-Based Latent Steering for Decoder-Only Time Series Forecasting

Time series forecasting aims to predict future values from historical observations and auxiliary features. We propose \textbf{SteerCast}, a retrieval-based latent steering method that improves decoder-only forecaster at inference time, without updating its parameters. SteerCast constructs a database from the training set by storing a representation of each history window together with a \emph{steering vector} computed in the forecaster's latent space, defined as the difference between representations induced by the ground-truth continuation and by the model's own prediction. At test time, SteerCast retrieves nearest neighbors for a query history, aggregates their steering vectors, and injects the resulting signal into the forecaster's hidden states at every step of autoregressive generation, guiding predictions toward trajectories consistent with similar training cases. Experiments across diverse multivariate benchmarks and multiple horizons show that SteerCast consistently improves forecasting accuracy over the fine-tuned backbone and retrieval-based baselines, while requiring no additional training beyond the original fine-tuning and using only the training set as a retrieval corpus.

cs.LG↗

CORAL: Cross-modal Vector Retrieval via Incremental Graph Construction at Scale

Cross-modal vector retrieval is widely used in multimodal systems, such as search engines and vector databases. It typically operates in out-of-distribution (OOD) settings, where query vectors follow a distribution that differs from that of the vectors stored in the database. In such cases, conventional indexes suffer significant performance degradation, and even methods specially designed for OOD remain limited by inefficient use of query modal characteristics, restricted GPU parallelism, and inadequate support for dynamic updates. We present CORAL, a novel GPU-accelerated graph-based vector index for scalable cross-modal retrieval, featuring hierarchical memory management that spans GPU, CPU, and disk. Specifically, CORAL incrementally incorporates the characteristics of query modality and terminates index construction timely. Crucially, it introduces coverage-aware adaptive pruning to address the imbalanced coverage of the query vector's neighbors. Moreover, CORAL presents a fully neighborhood-aware projection approach to efficiently utilize GPUs for highly parallel index construction, and a targeted connectivity enhancement method to refine the index structure. Besides, CORAL also supports modal-semantics-based vector insertion and topology-repairing deletion that restore node connectivity. Experimental results demonstrate that CORAL outperforms existing methods with up to 1.6 times the throughput at matched recall while reducing construction time by up to 56%. Furthermore, it exhibits remarkable resilience under dynamic updates and remains effective at the billion scale.

cs.DB↗

Harness Compilation: Which Decisions Should a Small Vision-Language Model Keep?

Small vision-language models may be able to read external evidence yet struggle to obtain it. We introduce Harness Compilation (HC), an offline procedure that adapts the division of work between a frozen small VLM and its external harness. A large teacher uses student execution traces to revise reusable content and control, while a separate validation set selects the deployed harness. Deployment requires neither weight updates nor teacher calls. Across seven visual question-answering settings with students of at most 9B parameters, HC improves scores over bare students by 9.9-23.9 points, averaged over three independent builds per setting. Interventions on five runtime decision types (invocation, selection, argument generation, evidence integration and abstention) show why this allocation matters: requesting evidence and generating open queries can be costly, whereas bounded choices and reading supplied text can remain useful student work. Fact cards benefit all ten evaluated students, but decision policies transfer unevenly. Recompilation for a new student model helps when the transferred interface no longer fits the student. With 100 practice items, HC exceeds answer-only LoRA on three tasks. Larger training budgets can match or surpass a fixed harness, while combining the two improves SlideVQA beyond either alone. These findings support allocating work from measured student behavior rather than uniformly removing decisions.

cs.AI↗

Localization of Candidate Kikuchi Regions in RHEED Images: Visibility and Annotation Boundaries

Kikuchi lines and bands in reflection high-energy electron diffraction (RHEED) carry information on crystal geometry and electron scattering, but are often obscured by intense diffraction streaks. We combine multiscale convolution with spatial detail from skip connections and independent supervision that allows overlapping regions to localize streaks and candidate Kikuchi regions separately. Using manual polygon annotations made before model-assisted editing, the Kikuchi intersection-over-union (IoU) on 29 laboratory test images from separate growth batches was $0.6290 \pm 0.0137$. We then performed supervised adaptation to public chalcogenide images and stratified images by Kikuchi visibility using the median ratings of three observers who rated the images separately, two of whom rated with predictions hidden. With all 244 training images, IoU for the clear-feature group of 10 test images was $0.5014 \pm 0.0336$, whereas ambiguous images gave lower values. With 100 training images, clear-group IoU was $0.5171 \pm 0.0073$; expanding the training set also reduced responses on images without target features. Reported uncertainties are sample standard deviations across four laboratory runs or three adaptation runs. Public-image adaptation used reference masks of mixed provenance, including prediction-derived drafts. The method converts visual cues into inspectable spatial regions, providing a basis for analysis of line positions, intersections, and local intensity.

cond-mat.mtrl-sci↗

MiniVer-V: Identifying Minimal Sufficient Evidence for Short Video Verification

A core challenge in short-video fact-checking is identifying which evidence is sufficient to support a verification conclusion. Existing approaches either give the verifier all available evidence, introducing noise, or select evidence by topical relevance, which conflates relatedness with sufficiency. We identify evidential sufficiency as the selection criterion: whether a subset of evidence is adequate to support a confident verdict without redundancy. We introduce MiniVer-V, a benchmark of 195 short videos with three-way verdict annotations (supported, refuted, insufficient) and 5,510 multimodal evidence units spanning visual keyframes, speech transcripts, and web-retrieved external sources. We propose a two-layer verification framework that separates claim-video consistency, assessed from internal evidence, from factual verdict determination, which additionally requires external corroboration. On top of it, a sufficiency-driven greedy search assembles evidence until a sufficiency threshold is met and outputs insufficient when the candidate pool is exhausted, rather than forcing a verdict. With Claude Sonnet 4, the method reaches a Macro-F1 of 0.510 using 4.5 evidence units on average (16% of the full evidence set), statistically indistinguishable from the full-evidence baseline (0.518 with 27.7 units), while significantly improving recognition of insufficient cases over the same search without abstention. The efficiency result replicates with GPT-5.5 and holds only partially with an open-weight Qwen2.5-72B verifier. Ablations show that external evidence is indispensable for factual determination, while internal video evidence grounds the verdict in claim-video consistency. These findings suggest that evidence-efficient verification is achievable, and that explicit abstention is needed when evidence is genuinely inadequate.

cs.CV↗

Exact polariton-condensate states via nonlinear Stark coupling

We study a hybrid system of coupled cavities, with or without two-level atoms doped inside. By employing a building-block construction facilitated by nonlinear Stark coupling, we identify a manifold of exact many-body eigenstates that remain analytically tractable across arbitrary lattice geometries and atom distributions. We demonstrate that these exact states exhibit a dual physical nature depending on their spectral positioning: when embedded within the chaotic many-body spectrum, they manifest as quantum scars that violate the eigenstate thermalization hypothesis and undergo weak ergodicity breaking; conversely, as ground states within fixed-total-excitation sectors, they exhibit off-diagonal long-range order and form polariton condensates. Furthermore, numerical calculations of local excitation-number fluctuations suggest that negative Stark coupling shifts the Mott-insulator to superfluid transition toward weaker inter-cavity hopping. Our findings provide a versatile framework for engineering both coherent quantum phases and non-thermalizing states in diverse light-matter architectures.

quant-ph↗

Markov length can diverge in systems whose universal physics is spatially Markovian

Conditional mutual information (CMI) probes spatial non-Markovianity and plays a central role in a recently developed framework for defining mixed-state phases based on local reversibility. This framework sharply distinguishes exponentially decaying from algebraically decaying CMI, since the latter obstructs equivalence to a state with finite Markov length. Here we show that states that flow to the same renormalization-group (RG) fixed point can nevertheless exhibit qualitatively different CMI. Starting from a critical Gibbs state of a finite-range local classical Hamiltonian, whose CMI vanishes beyond the interaction range, we show that an RG-irrelevant perturbation that breaks detailed balance generically generates algebraically decaying CMI. Crucially, these power-law tails are cutoff-suppressed: although they lead to an infinite Markov length at every fixed lattice cutoff, they vanish as a positive power of the cutoff in the continuum limit. We construct solvable models that exhibit this phenomenon and predict that the Ising-symmetric critical point of Toom's cellular automata generically exhibits such behavior. We also find analogous cutoff-suppressed CMI in the high-temperature paramagnetic phase of the long-range Ising paramagnet, in contrast with the genuine power-law CMI at the finite-temperature critical point in this same model that survives the continuum limit. We also derive a general result that in one spatial dimension, CMI decays faster than the inverse square of the buffer length if and only if the long-distance physics is Markovian.

cond-mat.stat-mech↗

Low-Cost Sensor Calibration for Indoor Air Quality Monitoring: A Dataset, Evaluation Scenarios, and a Lightweight Model

Low-cost sensors enable scalable indoor air quality monitoring but require calibration because of nonlinear distortions, noise, and temporal drift. The conventional strict pairwise calibration setting requires a co-located reference sensor at each deployment location and does not account for spatial and temporal heterogeneity. To address these limitations, we introduce a six-month dataset comprising multivariate indoor air-quality measurements from low-cost and reference sensors with contextual metadata collected at five locations. Using this dataset, we define four evaluation scenarios. The reference-efficient and location-transfer scenarios evaluate spatial generalization, whereas the long-term drift and event-conditioned scenarios assess robustness to gradual and abrupt distribution shifts. Based on these scenarios, we derive design requirements and propose a lightweight temporal model that combines input-window compression with residual temporal and feature fusion. Experiments show strong calibration performance across all four scenarios with low edge-inference cost.

cs.LG↗

Breaking the Group Size Barrier: Parameter-Efficient Group Dance Generation with Chain-of-Dancers

Group dance generation aims to synthesize coordinated multi-dancer choreography from music, with broad applications in animation and interactive content creation. This task requires modeling dense inter-person dependencies to ensure spatial coordination, while naturally preserving individual dancer identities. Existing approaches model all dancers jointly with end-to-end transformers, which tie the architecture to a fixed group size and entangle per-dancer identities across frames. We propose ChainDance, a scalable framework that reformulates group dance generation as a Chain-of-Dancers: a sequential decomposition over per-dancer conditional distributions, allowing a single model to scale across variable group sizes without retraining and naturally preserving per-dancer identity. Built on a frozen single-dancer diffusion backbone, ChainDance introduces two lightweight modules: a Role-Aware Text Encoder (RATE) for per-dancer semantic conditioning, and a Group-Aware Motion Encoder (GAME) that aggregates previously generated dancers via a distance-weighted graph convolutional network, and incorporates a training-free noise optimization procedure at inference time to enforce global spatial coherence. Experiments on AIOZ-GDance demonstrate that ChainDance achieves state-of-the-art motion quality and group coordination while structurally preserving per-dancer identity, with $3$-$4\times$ fewer parameters and requiring $3$-$6\times$ less training time compared to prior approaches.

cs.CV↗

Attributing HOW, Not Just WHICH: Counterfactual Response Trajectories for Diffusion Models

Diffusion models have achieved remarkable success in image generation, yet tracing their outputs to individual training examples remains challenging. Existing attribution methods often compress factor-specific effects into scalar responses, making distinct internal changes indistinguishable. This is particularly limiting for diffusion models, where semantic factors emerge through evolving representation dynamics during denoising. We therefore reformulate diffusion data attribution as attributing factor-induced internal response trajectories. In this paper, we propose a novel Concept Attribution method through Dynamic Trajectories(CADT). We argue that attribution should therefore ask not only \emph{which} examples matter, but also \emph{how} their influence unfolds during generation. Specifically, we construct matched counterfactual pairs at identical noisy states to isolate factor-specific representation displacements, and model their directional and magnitude evolution across denoising as dynamic attribution signatures. For each training example and generated query, CADT extracts stage-wise feature vectors and integrates them along the denoising process to form a trajectory descriptor. Applying the same construction across the training set yields a bank of factor-specific trajectory descriptors. The covariance statistics of this bank are then used to construct . CADT uses this covariance-aware positive-semidefinite kernel to calibrate the query and training representations, and compares the calibrated query trajectory with each training trajectory to produce the final training-sample attribution scores. Experiments on multiple public datasets show consistent improvements over existing diffusion attribution baselines across hierarchical, compositional, and style attribution.

cs.CV↗

Optimal upper tail estimates for the edge eigenvalues of $β$-ensembles

Hermite and Laguerre $β$-ensembles are important and well studied models in random matrix theory, with the special cases $β=1,2,4$ corresponding to classical matrix ensembles. Although sharp moderate-deviation estimates have been established for the largest eigenvalues, substantially less is known for other edge eigenvalues. Even for the limiting distributions $TW_β^{(k)}$, the leading constants in the tail asymptotics are not known for general $β$ and $k\ge2$. In this paper, for each fixed $k\ge2$, we prove matching upper and lower moderate deviation bounds for the right tail of the $k$-th largest eigenvalue, with the optimal exponential constant $2βk/3$, for general $β$. We also establish sharp lower tail estimates for the second largest eigenvalue of the Gaussian and Laguerre orthogonal ensembles, with exponential constant $1/24$, same as for the largest eigenvalue. For $TW_β^{(k)}$, we obtain matching right-tail asymptotics for $β\ge2/k$ and left-tail asymptotics for $0<β\le2$ for $k=2$. We also obtain new stochastic domination results for the edge eigenvalues. Our proofs combine variational arguments from tridiagonal matrix models, last passage percolation, stochastic comparisons of sums of edge eigenvalues, and superposition-decimation identities.

math.PR↗