Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 145 records · Page 8Linked to original sources

Kairos: Toward Adaptive and Parameter-Efficient Time Series Foundation Models

Inherent temporal heterogeneity, such as varying sampling densities and periodic structures, has posed substantial challenges in zero-shot generalization for Time Series Foundation Models (TSFMs). Existing TSFMs predominantly rely on massive parameterization to absorb such heterogeneity, as their static tokenization and positional encoding schemes entangle diverse temporal patterns into a fixed representation space, encouraging memorization rather than adaptation. To address this limitation, we propose Kairos, a flexible and parameter-efficient TSFM dedicated to forecasting tasks, which decouples temporal heterogeneity from model capacity through a novel tokenization perspective. Kairos introduces a dynamic patching tokenizer and a mixture-of-size encoding that adapt observational granularity to local information density, enabling fine-grained temporal abstraction without increasing model width or depth. In addition, we design a multi-granularity positional embedding based on dynamic rotary encodings, which conditions on instance-level spectral features and temporal structure induced by dynamic patching tokenization, allowing robust modeling of diverse temporal dependencies. Trained on a novel Predictability-Stratified Time-Series (PreSTS) corpus, Kairos achieves superior zero-shot performance with substantially fewer parameters on two mainstream benchmarks, GIFT-Eval and Time-Series-Library. The project page is at https://foundation-model-research.github.io/Kairos .

cs.LG↗

Edge GPU Aware Multiple AI Model Pipeline for Accelerated MRI Reconstruction and Analysis

Advancements in AI have greatly enhanced the medical imaging process, making it quicker to diagnose patients. However, very few have investigated the optimization of a multi-model system with hardware acceleration. As specialized edge devices emerge, the efficient use of their accelerators is becoming increasingly crucial. This paper proposes a hardware-accelerated method for simultaneous reconstruction and diagnosis of \ac{MRI} from \ac{CT} images. Real-time performance of achieving a throughput of nearly 150 frames per second was achieved by leveraging hardware engines available in modern NVIDIA edge GPU, along with scheduling techniques. This includes the GPU and the \ac{DLA} available in both Jetson AGX Xavier and Jetson AGX Orin, which were considered in this paper. The hardware allocation of different layers of the multiple AI models was done in such a way that the ideal time between the hardware engines is reduced. In addition, the AI models corresponding to the \ac{GAN} model were fine-tuned in such a way that no fallback execution into the GPU engine is required without compromising accuracy. Indeed, the accuracy corresponding to the fine-tuned edge GPU-aware AI models exhibited an accuracy enhancement of 5\%. A further hardware allocation of two fine-tuned GPU-aware GAN models proves they can double the performance over the original model, leveraging adequate partitioning on the NVIDIA Jetson AGX Xavier and Orin devices. The results prove the effectiveness of employing hardware-aware models in parallel for medical image analysis and diagnosis.

cs.AR↗

Demystifying MaskGIT Sampler and Beyond: Adaptive Order Selection in Masked Diffusion

Masked diffusion models have shown promising performance in generating high-quality samples in a wide range of domains, but accelerating their sampling process remains relatively underexplored. To investigate efficient samplers for masked diffusion, this paper theoretically analyzes the MaskGIT sampler for image modeling, revealing its implicit temperature sampling mechanism. Through this analysis, we show that MaskGIT is asymptotically equivalent to a choose-then-sample (CTS) formulation, instantiated as the "moment sampler," which explicitly separates index selection from token sampling. This CTS reformulation is essential: it yields unbiased token sampling and exposes an algorithmic design space for index selection, both of which are inaccessible in MaskGIT's original formulation. Regarding token sampling, we reveal that MaskGIT implicitly adopts a low-temperature sampler, which explains why MaskGIT often degrades with more sampling steps. The CTS reformulation of MaskGIT allows us to fix the temperature sampling to ensure unbiasedness. We also improve the index selection in CTS through two key innovations: a partial caching technique for transformers that approximates longer sampling trajectories without proportional computational cost, and a hybrid approach formalizing the exploration-exploitation trade-off in adaptive unmasking. Experiments in image and text domains demonstrate our theory as well as the efficiency of our proposed methods, advancing both theoretical understanding and practical implementation of masked diffusion samplers.

cs.LG↗

What your brain activity says about you: A review of neuropsychiatric disorders identified in resting-state and sleep EEG data

Electroencephalogram monitoring devices and online data repositories hold large amounts of data from individuals participating in research and medical studies without direct reference to personal identifiers. This paper explores what types of personal and health information have been detected and classified within task-free EEG data. Additionally, we investigate key characteristics of the collected resting-state and sleep data, in order to determine the privacy risks involved with openly available EEG data. We used Google Scholar, Web of Science and searched relevant journals to find studies which classified or detected the presence of various disorders and personal information in resting state and sleep EEG. Only English full-text peer-reviewed journal articles or conference papers about classifying the presence of medical disorders between individuals were included. A quality analysis carried out by 3 reviewers determined general paper quality based on specified evaluation criteria. In resting state EEG, various disorders including Autism Spectrum Disorder, Parkinson's disease, and alcohol use disorder have been classified with high classification accuracy, often requiring only 5 mins of data or less. Sleep EEG tends to hold classifiable information about sleep disorders such as sleep apnea, insomnia, and REM sleep disorder, but usually involve longer recordings or data from multiple sleep stages. Many classification methods are still developing but even today, access to a person's EEG can reveal sensitive personal health information. With an increasing ability of machine learning methods to re-identify individuals from their EEG data, this review demonstrates the importance of anonymization, and the development of improved tools for keeping study participants and medical EEG users' privacy safe.

cs.NE↗

Stable Central Limit Theorems for Discrete-Time Lag Martingale Difference Arrays: Applications to Dynamic Causal Inference

Recent work in dynamic causal inference introduced a class of discrete-time stochastic processes that generalize martingale difference sequences and arrays as follows: the random variates in each sequence have expectation zero given certain lagged filtrations but not given the natural filtration. We formalize this class of stochastic processes and prove stable central limit theorems (CLTs) via martingale-coboundary decomposition, leveraging the classical martingale CLT. We develop a variety of sufficient conditions, including conditions under which the limiting variance has a simple form that depends on variances and covariances of neighboring variates. We demonstrate the application of these results to inference for time-averaged treatment effects in switchback designs and present a simulation study supporting their validity. The CLTs enable various extensions to existing methodology for design-based approaches to dynamic causal inference, including time-lagged effects, random limiting variances, cross-unit dependence, and vector-valued estimands.

math.ST↗

Large Language Model Selection with Limited Annotations

Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotations over fixed evaluation sets. To address this challenge, we develop SELECT-LLM, the first framework for active model selection of LLMs. SELECT-LLM aims to find a small set of queries whose annotations are most informative for identifying the best LLM for a given task. To this end, we introduce a query selection rule based on expected information gain, computed from pairwise similarities between candidate model outputs. Because this rule only uses generated model responses, SELECT-LLM can be applied across candidate models without assumptions about their architecture or access to model weights. This makes it suitable for both open-weight and black-box LLMs. We evaluate SELECT-LLM across 23 datasets, 156 evaluated models, diverse task families, and multiple text evaluation metrics. Across all experiments, SELECT-LLM improves over the strongest baseline in every setting, with annotation cost reductions up to 81.8% for best model selection and up to 84.78% for near-best model selection.

cs.CL↗

DE-Sinc approximation for unilateral rapidly decreasing functions and its computable error bound

The Sinc approximation is highly effective for functions that decay rapidly at both ends of the real axis. For unilateral rapidly decreasing functions, which decay algebraically as $t\to-\infty$ and exponentially as $t\to\infty$, an appropriate variable transformation is required. Existing single-exponential transformations for this class of functions yield only root-exponential convergence, even when improved transformations are employed. This paper develops a double-exponential (DE)-Sinc approximation based on the transformation $t=2\sinh(\log(\log(1+\exp(π\sinh x))))$, which was previously introduced for numerical integration. The main contribution is a rigorous and computable error bound of order $\operatorname{O}(\exp(-cn/\log n))$ with a constant explicitly expressed in terms of the problem parameters. Under the stated assumptions, the resulting approximation achieves almost exponential convergence and is suitable for computation with guaranteed accuracy. Numerical examples satisfying these assumptions confirm the predicted convergence behavior and the validity of the derived error bound.

math.NA↗

NOSA: Native and Offloadable Sparse Attention

Decoding throughput improvements from larger inference batches are limited by GPU memory, which is largely consumed by the key-value (KV) cache. Prior training-free KV cache offloading alleviates this by keeping redundant context on the CPU and fetching only a sparse subset for attention, but it often degrades long-generation quality due to training-inference mismatch on sparse patterns. Meanwhile, trainable sparse attention is incompatible with efficient offloading, as unconstrained KV accesses may force large CPU-to-GPU transfers and erase throughput gains. To this end, we propose NOSA, a trainable sparse attention mechanism natively designed for KV cache offloading. NOSA explicitly constrains the volume of CPU-GPU KV transfers, thereby achieving low communication overhead and high decoding throughput. We further build NOSI, a KV cache offloading inference system that fully unlocks NOSA's efficiency. Empirical results on 1,3,8B LLMs demonstrate that NOSA outperforms KV cache offloading baselines on general, long-input, and long-generation tasks, while boosting decoding throughput by up to 5.04x, 1.92x, and 1.83x over FullAttn, InfLLMv2, and ShadowKV, respectively. We release our code at https://github.com/thunlp/NOSA.

cs.CL↗

Scaling Vision Transformers for Functional MRI with Flat Maps

We study the problem of training self-supervised foundation models for functional MRI. Our main contributions are: (1) we introduce a new model family (CortexMAE) trained using the masked autoencoder framework on 2.1K hours of open fMRI data, and (2) we release the first open evaluation suite (Brainmarks) for fMRI foundation models. Our core innovation is simple: we adapt the Vision Transformer to fMRI by first converting each 3D fMRI volume to a 2D map using a cortical flat map projection. We directly compare flat maps to both parcellation and volume-based representations. While each has its advantages, flat maps generally perform best. We perform the first systematic scaling analysis for fMRI and observe strict power law scaling, albeit with limits. Finally, we use Brainmarks to do controlled benchmark comparisons. On subject-level trait prediction, we report a challenging null result: no single model achieves clear state-of-the-art performance. Moreover, all models struggle to outperform a simple functional connectivity baseline. On cognitive state decoding, we observe more robust performance, and in this setting our CortexMAE family outperforms prior models by a large margin. Code, models, and datasets are available at https://github.com/MedARC-AI/CortexMAE and https://github.com/MedARC-AI/Brainmarks.

cs.CV↗

Stable toric sheaves. I : Chern classes

We reduce Hartshorne's conjecture on indecomposable rank 2 vector bundles in the semistable case to the non-existence of smoothable rank 2 semistable toric sheaves on the projective space. We then study Chern classes for rank 2 torsion-free toric sheaves. For the reflexive ones, we derive a simple formula for the Chern polynomial, and in the general torsion-free case we introduce an iterative construction method based on elementary injections, allowing us to prescribe Chern classes. This yields infinite families of explicit examples on P4 and P5, and establishes existence on Pn for all n greater than 3, with Chern classes satisfying all known constraints arising from locally freeness and indecomposability. We also provide simple obstructions for smoothability.

math.AG↗

Theoretical Refinement of CLIP by Utilizing Linear Structure of Optimal Similarity

In this study, we propose an enhancement to the similarity computation mechanism in multimodal contrastive pretraining frameworks such as CLIP. Prior theoretical research has demonstrated that the optimal similarity metrics between paired modalities should correspond to the pointwise mutual information (PMI) between the two modalities. However, the current implementations of CLIP and its variants fail to fully utilize the underlying linear structure of PMI. We therefore propose KME-CLIP, which leverages this structure through the inner product in a reproducing kernel Hilbert space (RKHS). We theoretically prove that, under our assumptions, the KME-CLIP similarity can bring the contrastive loss arbitrarily close to its optimal value, which is attained by PMI, as the size of the point set grows, and we empirically evaluate KME-CLIP against CLIP and its kernel-based variants across several retrieval and classification tasks.

cs.LG↗

Fluctuation-Response Theory of Non-Equilibrium Complex Fluids

A fundamental challenge in soft matter physics is to describe materials, such as the living cytoplasm and tissues, that are simultaneously active, chemically driven, and exhibit long-lasting memory of mechanical stresses. Here, we construct a generalized hydrodynamic framework at finite wavevectors and frequencies that is applicable to non-equilibrium fluids with memory. By leveraging stationary correlation identities, we derive a generalized linear response theory for non-equilibrium steady states. This framework serves as a formal extension of Onsager's regression hypothesis beyond thermal equilibrium. Our approach provides a direct pathway to derive transport coefficients from steady-state fluctuations without the traditional Mori-Zwanzig projection-operator formalism, generalizing the Green-Kubo relations to non-equilibrium systems. As a corollary, we derive two model-free variants of the non-equilibrium fluctuation-response relation for non-Markovian dynamics. These generalized relations explicitly capture the non-equilibrium circulating currents--an out-of-equilibrium signature that is invisible to conventional scalar formulations or frameworks that treat degrees of freedom independently. Applying our theory to chemically driven active fluids reveals the emergence of active viscoelastic memory, wherein chemical reaction cycles dynamically renormalize the macroscopic viscous response. Strikingly, this active memory can induce negative storage and loss moduli at finite frequencies, a behavior absent in ordinary viscoelastic fluids. Our first-principles framework rigorously extends linear rheology to non-equilibrium systems and provides a foundation for understanding non-Markovian dynamics across a broad range of biological and synthetic active matter.

cond-mat.soft↗

HumanHalo: Safe and Efficient 3D Navigation Among Humans via Minimally Conservative MPC

Safe and efficient robotic navigation among humans is essential for integrating robots into everyday environments. Most existing approaches focus on simplified 2D crowd navigation and fail to account for the full complexity of human body dynamics beyond root motion. We present HumanHalo, an MPC framework for 3D MAV navigation among humans that combines theoretical safety guarantees with data-driven models for realistic human motion forecasting. Our approach introduces a novel reachability-based safety formulation that constrains only the initial control input for safety while modeling its effects over the entire planning horizon, enabling safe yet efficient navigation. We validate HumanHalo in both simulated experiments using real human trajectories and in the real world, demonstrating its effectiveness across tasks ranging from goal-directed navigation to visual servoing for human tracking. While we apply our method to MAV in this work, it is generic and can be adapted to other platforms. Our results show that the method preserves safety without excessive conservatism, even under imperfect human prediction and in constrained workspaces, while remaining efficient enough for real-time onboard execution.

cs.RO↗

Constrained Ramsey numbers for rainbow $P_5$

For any graph $H$, let $R_k(H)$ denote the minimum integer $n$ such that in every $k$-edge-coloring of $K_{n}$, there is a monochromatic copy of $H$. Given two graphs $H$ and $G$, the \emph{constrained Ramsey number} $f(H,G)$ is the minimum integer $n$ such that in every edge-coloring of $K_{n}$ with any number of colors, there is either a monochromatic copy of $H$ or a rainbow copy of $G$. Let $P_t$ be the path on $t$ vertices. Gyárfás, Lehel and Schelp proved that $f(H,P_5)=R_3(H)$ when $H$ is a path, a cycle, or a connected non-bipartite graph. Li, Besse, Magnant, Wang and Watts conjectured that $f(H,P_5)=R_3(H)$ for any graph $H$, and confirmed this for every connected graph and every bipartite graph. In this paper, we address this conjecture for several classes of disconnected graphs with chromatic number at least 3. To the best of our knowledge, our general results encompass all previously known results of this type. We also obtain several results for a bipartite variant of the problem. In addition, we propose a series of related questions from several directions for further research. Our proofs combine the structural characterization of Thomason and Wagner with Simonovits' decomposition-family method and Ramsey numbers for families of graphs.

math.CO↗

TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications

Large Language Models (LLMs) are increasingly deployed in complex multi-agent applications that rely on external function calls. This workload creates severe performance challenges for the KV Cache: spatial contention leads to the eviction of critical agents' caches and temporal underutilization leaves the cache of agents stalled on long-running function calls idling in GPU memory. We present TokenCake, a KV-Cache-centric serving framework that bridges this gap by co-optimizing scheduling and memory management through an agent-aware design. TokenCake's Temporal Scheduler employs an event-driven, opportunistic policy to proactively offload idle KV Caches during function calls and uses predictive uploading to hide data transfer latency. TokenCake's Spatial Scheduler uses dynamic memory partitioning, guided by a hybrid priority metric combining graph structure and runtime state, to reserve GPU memory for critical-path agents. Our evaluation on representative multi-agent benchmarks shows that TokenCake reduces end-to-end latency by up to 47.06% in memory-constrained settings and improves effective GPU utilization by 16.9 percentage points compared to vLLM. TokenCake is publicly available at https://github.com/pkulemonade/TokenCake.

cs.DC↗

Managing Self-Learning Experts under Per-Round Budget Constraints

This paper addresses the problem of sequential decision-making under learning budget constraints. Such settings naturally arise in applications like managing a portfolio of bandit or reinforcement learning (RL) algorithms. We propose a novel UCB-type algorithm, M-LCB, designed to manage a pool of $K$ self-learning experts in a stochastic environment while accounting for a limited per-round learning budget $M$. At each round, M-LCB selects one expert to make a decision and at most $M \le K$ experts to learn. For selection, M-LCB uses confidence bounds constructed from limited prior knowledge about the experts (i.e., mild assumptions) and their observed training losses. We derive anytime regret bounds for M-LCB that scale with the individual regrets of the experts. In particular, if each expert has regret $\tilde O(T^α)$ by round $T$, then M-LCB guarantees an overall regret of $\tilde O\left(\sqrt{KT/M} + (K/M)^{1-α}T^α\right)$ relative to the best expert in hindsight. Finally, we demonstrate the applicability of M-LCB using self-learning experts instantiated as (i) parametric models and (ii) bandit algorithms.

cs.LG↗

Ultrafast Processing of Hyper-Entangled Bell States at the Optical Bandwidth Limit

Bell states embody the most basic form of two-state entanglement and are a key component of quantum protocols of communication and sensing, indicating that high-speed processing of polarization Bell-states is highly desired for quantum technology. Yet, the processing speed of all current protocols is inherently limited by the electronic bandwidth of photo-detectors (Photomultiplier tubes, avalanche photo-diodes, etc.) that can process only $10^{6-7}$ photons/s, whereas standard broadband sources may easily produce $10^{10-14}$ photons/s. Using dual polarization, nonlinear SU(1,1) interferometry as an ultrafast quantum detector of the entangled bi-photons, we demonstrate the complete cycle of processing - generation, manipulation \textit{and detection} of polarization-entangled Bell states at an ultra-high photon-flux of $\sim\!5\times10^{11}$ photons/s ($\sim\!5$ orders of magnitude higher than standard methods), limited only by the optical bandwidth of the generating source.

quant-ph↗

A Survey on Efficient Vision-Language-Action Models

Vision-Language-Action models (VLAs) represent a significant frontier in embodied intelligence, aiming to bridge digital knowledge with physical-world interaction. Despite their remarkable performance, foundational VLAs are hindered by the prohibitive computational and data demands inherent to their large-scale architectures. To this end, recent studies improve VLA efficiency from different views, e.g., real-time inference, training computation, and scalable data collection. However, these efforts are mostly studied separately. A unified view is still missing for understanding how efficiency should be optimized across the full VLA lifecycle. To bridge this gap, this survey presents the first comprehensive review of Efficient Vision-Language-Action models (Efficient VLAs) across the entire model-training-data pipeline. Specifically, we introduce a unified taxonomy to systematically organize the disparate efforts in this domain, categorizing current techniques into three core pillars: (1) Efficient Model Design, focusing on efficient architectures and model compression; (2) Efficient Training, which reduces computational burdens during model learning; and (3) Efficient Data Collection, which addresses the bottlenecks in acquiring and utilizing robotic data. Through a critical review of state-of-the-art methods within this framework, this survey provides an organized reference for the community and summarizes representative applications, delineates key challenges, and charts a roadmap for future research. We maintain a continuously updated project page to track our latest developments: https://evla-survey.github.io/.

cs.CV↗