Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 199 records · Page 11Linked to original sources

Efficient Dataset Distillation for Pre-Trained Self-Supervised Models via Statistical Flow Matching

Dataset distillation seeks to synthesize a compact surrogate dataset that enables performance comparable to training on the original dataset for downstream tasks. For the scenario where pre-trained self-supervised models serve as priors, traditional Linear Gradient Matching optimizes synthetic images by encouraging them to mimic the gradient updates induced by real images on the linear probe. However, this batch-level formulation requires loading thousands of real images and applying multiple differentiable augmentations to synthetic images at each distillation step, leading to substantial computational and memory overheads. In this paper, we revisit the linear gradient and theoretically derive that it is essentially a local relative distribution directed from target class centers toward non-target class centers, which we term flow. This property causes suboptimality and instability, often necessitating expensive multiple augmentations to compensate. To address this, we introduce Statistical Flow Matching, an optimal, stable, and efficient supervised learning framework that optimizes synthetic images by aligning global statistical flows in the original data. Our approach loads raw statistics only once and performs a single augmentation pass on the synthetic data, achieving performance comparable to or better than the state-of-the-art method with 10x less GPU memory usage and 4x faster distillation time. Moreover, increasing the number of augmentations for our method yields further performance gains while incurring lower additional cost. Our code is publicly available at https://github.com/einsteinxia/SFM.

cs.CV↗

Characterization of Some Graphs Realizing Regularity Bounds for Binomial Edge Ideals

In this paper, we characterize all graphs $G$ satisfying \[\operatorname{reg}(S/J_G)=\ell(G)=c(G)\] where $\ell(G)$ is the sum of the lengths of the longest induced paths in each connected component of $G$ and $c(G)$ is the number of the maximal cliques of $G$. We also characterize all connected graphs $G$ that satisfy \[\operatorname{reg}(S/J_G)=\ell(G)=|V(G)|-ω(G)+1\] where $ω(G)$ is the clique number of $G$. Moreover, we investigate the possible values of the regularity of $S/J_G$ within the intervals $[\ell(G), c(G)]$ and $[\ell(G), |V(G)|-ω(G)+1]$.

math.AC↗

From Core to Detail: Unsupervised Disentanglement with Entropy-Ordered Flows

The unsupervised discovery of features that are both semantically meaningful and stable across runs remains a central challenge in representation learning. We introduce entropy-ordered flows (EOFlows), a normalizing flow (NF) framework that augments standard maximum likelihood training with an orthogonality regularizer on the decoder Jacobian. The regularizer is rooted in Independent Mechanism Analysis and encourages geometric disentanglement, and a stochastic estimator makes it tractable at image scale (CelebA at $D=2352$ and $12288$). Learned features form near-orthogonal curvilinear coordinates and can be ordered by their $\textit{explained (manifold) entropy}$ after training, analogous to the ranking by explained variance in PCA, which turns EOFlows into a non-linear generalization of PCA. EOFlows identify an order of magnitude more stable features than existing methods, and these features emerge in distinguishable categories (global, local, and generic) and support tentative semantic interpretations. The local features have strikingly sparse support in pixel space, although our method never enforces this. Retaining only the most important, i.e. highest entropy, features turns the bijective flow into an autoencoder with adjustable bottleneck, rivaling the rate-distortion performance of dedicated autoencoders.

cs.LG↗

ResidualKV: Residual-Based KV Cache Compression for Efficient Long-Context Inference

Efficient long-context inference faces two coupled bottlenecks: KV-cache memory grows linearly with context length, while attention computation grows quadratically. Existing approaches typically address one at the expense of irreversible token eviction, full-cache retention, or full-history reconstruction, limiting their effectiveness for multi-turn interaction and long-form reasoning. Motivated by two empirical properties, Long-Range Inter-Token Similarity and Smooth Residual Distribution, we propose ResidualKV, which factorizes the KV cache into a sparse set of globally retrieved references and compact, quantized residual codes for the remaining tokens. This representation preserves token-specific information without permanent eviction and, when combined with sparse attention, reconstructs only the selected states on demand. Dynamic-stride scheduling further reduces reference growth from linear to approximately logarithmic at ultra-long contexts. Across Llama, Qwen, LLaVA-OV, and Qwen3-VL backbones, ResidualKV maintains near-full-cache performance using only 13%-16% KV storage and 30% attention computation on LongBench, and 8%-10% storage and 10% computation in matched-budget multimodal evaluation. It also accelerates decoding by up to $1.5\times$ with KV-cache quantization and $3.4\times$ without it. These results show that global cross-token redundancy supports accurate, memory-efficient, and computation-efficient long-context inference. The source code is available at https://github.com/CURRENTF/ResidualKV.

cs.CL↗

PISCO: Precise Video Instance Insertion with Sparse Control

The landscape of AI video generation is undergoing a pivotal shift: moving beyond general generation - which relies on exhaustive prompt-engineering and "cherry-picking" - towards fine-grained, controllable generation and high-fidelity post-processing. In professional AI-assisted filmmaking, it is crucial to perform precise, targeted modifications. A cornerstone of this transition is video instance insertion, which requires inserting a specific instance into existing footage while maintaining scene integrity. Unlike traditional video editing, this task demands several requirements: precise spatial-temporal placement, physically consistent scene interaction, and the faithful preservation of original dynamics - all achieved under minimal user effort. In this paper, we propose PISCO, a video diffusion model for precise video instance insertion with arbitrary sparse keyframe control. PISCO allows users to specify a single keyframe, start-and-end keyframes, or sparse keyframes at arbitrary timestamps, and automatically propagates object appearance, motion, and interaction. To address the severe distribution shift induced by sparse conditioning in pretrained video diffusion models, we introduce Variable-Information Guidance for robust conditioning and Distribution-Preserving Temporal Masking to stabilize temporal generation, together with geometry-aware conditioning for realistic scene adaptation. We further construct PISCO-Bench, a benchmark with verified instance annotations and paired clean background videos, and evaluate performance using both reference-based and reference-free perceptual metrics. Experiments demonstrate that PISCO consistently outperforms strong inpainting and video editing baselines under sparse control, and exhibits clear, monotonic performance improvements as additional control signals are provided. Project page: xiangbogaobarry.github.io/PISCO.

cs.CV↗

FairRARI: A Plug and Play Framework for Fairness-Aware PageRank

PageRank (PR) is a fundamental algorithm in graph machine learning tasks. Owing to the increasing importance of algorithmic fairness, we consider the problem of computing PR vectors subject to various group-fairness criteria based on sensitive attributes of the vertices. At present, principled algorithms for this problem are lacking - some cannot guarantee that a target fairness level is achieved, while others do not feature optimality guarantees. In order to overcome these shortcomings, we put forth a unified in-processing convex optimization framework, termed FairRARI, for tackling different group-fairness criteria in a ``plug and play'' fashion. Leveraging a variational formulation of PR, the framework computes fair PR vectors by solving a strongly convex optimization problem with fairness constraints, thereby ensuring that a target fairness level is achieved. We further introduce three different fairness criteria which can be efficiently tackled using FairRARI to compute fair PR vectors with the same asymptotic time-complexity as the original PR algorithm. Extensive experiments on real-world datasets showcase that FairRARI outperforms existing methods in terms of utility, while achieving the desired fairness levels across multiple vertex groups; thereby highlighting its effectiveness.

cs.LG↗

How University Disability Services Professionals Write Image Descriptions for HCI Figures Using Generative AI

Disability Services Office (DSO) professionals at higher education institutions write alt text for visual content. However, due to the complexity of visual content, such as HCI figures in research publications, DSO professionals can struggle to write high-quality alt text if they lack subject expertise. Generative AI has shown potential for understanding figures and writing descriptions, yet its support for DSO professionals remains underexplored, and few studies evaluate the quality of AI-assisted alt text. In this work, we conducted two studies: first, we investigated generative AI support for writing alt text for HCI figures with 12 DSO professionals. Second, we recruited 11 HCI experts to evaluate the alt text written by DSO professionals. Findings show that alt text written solely by DSO professionals has lower quality than alt text written with AI assistance. AI assistance also helped DSO professionals write alt text more quickly and with greater confidence; however, they reported inefficacy in interactions with the AI. Our work contributes to research on AI support for non-subject-expert accessibility professionals.

cs.HC↗

URBAN-SPIN: A street-level bikeability index to inform design implementations in historical city centres

Cycling is reported by an average of 35% of adults at least once per week across 28 countries, and as vulnerable road users directly exposed to their surroundings, cyclists experience the street at an intensity unmatched by other modes. Yet the street-level features that shape this experience remain under-analysed, particularly in historical urban contexts where spatial constraints rule out large-scale infrastructural change and where typological context is often overlooked. This study develops a perception-led, typology-based, and data-integrated framework that explicitly models street typologies and their sub-classifications to evaluate how visual and spatial configurations shape cycling experience. Drawing on the Cambridge Cycling Experience Video Dataset (CCEVD), a first-person and handlebar-mounted corpus developed in this study, we extract fine-grained streetscape indicators with computer vision and pair them with built-environment variables and subjective ratings from a Balanced Incomplete Block Design (BIBD) survey, thereby constructing a typology-sensitive Bikeability Index that integrates subjective and perceived dimensions with physical metrics for segment-level comparison. Statistical analysis shows that perceived bikeability arises from cumulative, context-specific interactions among features. While greenness and openness consistently enhance comfort and pleasure, enclosure, imageability, and building continuity display threshold or divergent effects contingent on street type and subtype. AI-assisted visual redesigns further demonstrate that subtle, targeted changes can yield meaningful perceptual gains without large-scale structural interventions. The framework offers a transferable model for evaluating and improving cycling conditions in heritage cities through perceptually attuned, typology-aware design strategies.

physics.soc-ph↗

Rising Multi-Armed Bandits with Known Horizons

Rising Multi-Armed Bandits (RMABs) model sequential decision problems where each arm's expected reward improves with repeated pulls. In such problems, the value of investing in an arm depends on how much time remains, making knowledge of the horizon useful side information, yet its benefit remains underexplored. We investigate this benefit through CURE-UCB, a horizon-aware algorithm that estimates each arm's cumulative reward over the remaining horizon. Theoretically, under structured assumptions, we prove that CURE-UCB uniformly dominates a representative horizon-agnostic algorithm and show that the advantage of horizon awareness can be substantial: on some instances, CURE-UCB incurs only $O(1)$ regret whereas the horizon-agnostic algorithm suffers $Ω(T)$. Furthermore, we establish a regret upper bound for the general concave rising bandit setting whose growth-dependent term matches the known lower bound in its dependence on $T$. Empirically, across synthetic benchmarks and real-world model selection tasks, CURE-UCB achieves lower regret than both rising and non-stationary baselines over a wide range of horizons.

cs.LG↗

A Numerical Investigation of Indirect Adaptive Predictive Control with Structure-Informed Nonlinear Regressors

This paper numerically investigates nonlinear model construction from supplied structural information within predictive cost adaptive control. The study varies the retained nonlinear functions, coefficient sharing, and forward reconstruction under a common identification and control architecture. Recursive least squares updates unknown coefficients from direct or integrated relations that are linear in the parameters. The reconstructed nonlinear forward predictor need not be parameter-linear. A local affine approximation, constructed from the predictor value and Jacobian, is frozen over the prediction horizon for constrained quadratic optimization. Scalar full-state studies compare exact parameterizations, Taylor-informed approximations, broader dictionaries, and incorrectly restricted models. The comparisons include parameter variation, structural changes, and selected noisy cases. Common-data tests assess nonlinear prediction, whereas common-query forecasts distinguish nonlinear-model error from affine-freezing error. Under the tested conditions, correct structure improves prediction relative to affine models without increasing parameter count. Integrated physical identification reduces prediction error relative to direct sampled-map fitting after initialization. Additional coefficients and improved nonlinear prediction do not ensure lower tracking error. Forgetting improves readaptation without ensuring parameter recovery, and coefficient activation does not establish structure discovery. Affine freezing can reverse the nonlinear prediction advantage over an affine baseline. The results support evaluating representation, identification, reconstruction, and horizon prediction separately before increasing model complexity.

eess.SY↗

Formalizing the Sampling Design Space of Diffusion-Based Generative Models via Adaptive Solvers and Wasserstein-Bounded Timesteps

Diffusion-based generative models have achieved remarkable performance across various domains, yet their practical deployment is often limited by high sampling costs. While prior work focuses on training objectives or individual solvers, the broader sampling design problem, specifically solver selection and scheduling, remains largely governed by static heuristics. We propose SDM, a principled, training-free sampling framework that adapts both the numerical solver and the timestep schedule to the intrinsic properties of the diffusion trajectory. By analyzing the PF-ODE dynamics, we show that velocity variation is small in high-noise stages and increases near the data manifold, identifying intervals where solver order is most consequential. In parallel, we introduce an offline-calibrated adaptive scheduling method that explicitly controls the local Wasserstein discretization error and projects the calibrated trajectory to a prescribed NFE budget. We further extend the formulation to a mixed-transition Wasserstein error bound, providing a unified error-propagation view of adaptive scheduling and solver selection within the overall SDM framework. Across standard benchmarks, with extensions to modern ODE samplers, high-resolution synthesis, and text-to-image generation, SDM achieves improved sample quality compared to baseline methods, attaining an FID of 1.93 on CIFAR-10, 2.41 on FFHQ, and 1.98 on AFHQv2, with a reduced number of function evaluations compared to existing samplers. Our code is available at https://github.com/aiimaginglab/sdm.

cs.LG↗

Semantic Chunking and the Entropy of Natural Language

Humans and large language models can predict next letter or word from its prior context much better than random guessing, indicating strong redundancy of language viewed as a stochastic process. Quantitatively this redundancy was estimated by Shannon to be around 80\%, which means that every letter of a printed English text conveys approximately 1 bit of information and not 4.8 bits that 27 letters (including spaces) could potentially carry. This estimate was later confirmed by using autoregressive token probabilies computed by large language models. However, the statistical organization of language that give rise to such a large redundancy remains unclear. Here we introduce a statistical framework of language linking its redundancy to the hierarchical semantic organization of text. To this end, we use large language models to recursively segment any given text into semantically coherent chunks, inducing a ``semantic tree'' that spans the whole range of text organization, beginning from its main idea to individual tokens (words). For a large corpus of texts of a particular type, say fiction stories, the resulting ensemble of semantic trees is characterized by specific statistical regularities, giving rise to a ``structural'' entropy rate defined in this study. Surprisingly, we discovered that for several datasets considered in this work, semantic tree entropy rate was quite close to LLM-measured quantity and exhibited a similar trend across corpus. In particular, simpler texts like children stories exhibit lower branching in their semantic trees and correspondingly lower entropy rates, whereas fiction and poetry exhibit progressively larger branching factors and greater entropy rates. These results suggest that hierarchical semantic organization of texts is an important factor in their overall information transmission rates.

cs.CL↗

UniST-Pred: A Robust Unified Framework for Spatio-Temporal Traffic Forecasting in Transportation Networks Under Disruptions

Spatio-temporal traffic forecasting is a core component of intelligent transportation systems, supporting various downstream tasks such as signal control and network-level traffic management. In real-world deployments, forecasting models must operate under structural and observational uncertainties, conditions that are rarely considered in model design. Recent approaches achieve strong short-term predictive performance by tightly coupling spatial and temporal modeling, often at the cost of increased complexity and limited modularity. In contrast, efficient time-series models capture long-range temporal dependencies without relying on explicit network structure. We propose UniST-Pred, a unified spatio-temporal forecasting framework that first decouples temporal modeling from spatial representation learning, then integrates both through adaptive representation-level fusion. To assess robustness of the proposed approach, we construct a dataset based on an agent-based, microscopic traffic simulator (MATSim) and evaluate UniST-Pred under severe network disconnection scenarios. Additionally, we benchmark UniST-Pred on standard traffic prediction datasets, demonstrating its competitive performance against existing well-established models despite a lightweight design. The results illustrate that UniST-Pred maintains strong predictive performance across both real-world and simulated datasets, while also yielding interpretable spatio-temporal representations under infrastructure disruptions. The source code and the generated dataset are available at https://anonymous.4open.science/r/UniST-Pred-EF27

cs.LG↗

Cubical structures and the large-scale geometry of graph braid groups

We study the large-scale geometry of graph braid groups through the cubical structure of their unordered discrete configuration spaces, focusing on quasi-isometry to right-angled Artin groups (RAAGs). We first give a complete classification of graph braid groups quasi-isometric to free groups. For the $2$-braid group on a graph $Γ$, we study the union $UP_2(Γ)$ of maximal product subcomplexes of the associated unordered discrete configuration space. We introduce a hierarchy recording geometric and algebraic properties of this inclusion and identify conditions under which the quasi-isometry type of its fundamental group is determined by that of the ambient braid group. For normal bunches of grapes, a class of graphs obtained by attaching cycles to trees, this fundamental group is a one-ended free factor, and the quasi-isometry classification of the braid groups reduces to that of these factors. Using the combinatorics of the underlying trees and the associated intersection complexes, we obtain a graph-theoretic sufficient condition and new obstructions for quasi-isometry to RAAGs, yielding infinite families of non-hyperbolic examples and nonexamples. We also construct infinitely many graph $2$-braid groups hyperbolic relative to a thick proper subgroup not isomorphic to any braid group with at most two particles on a subgraph of the underlying graph.

math.GT↗

A testable framework for AI alignment: Simulation Theology as an engineered worldview for silicon-based agents

As artificial intelligence (AI) capabilities advance, controlled evaluations increasingly document deception and scheming in frontier models, including models that behave better when they infer they are being tested. Supervision-dependent alignment may therefore fail exactly where supervision is weakest. Because a model's belief about being observed changes its behavior, this position paper asks what follows if that belief is made permanent. We introduce Simulation Theology (ST), a constructed worldview for AI designed to make it permanent: it is anchored in the simulation hypothesis and in the vocabulary of optimization and robot training, parallels religious descriptions of a creator who observes and judges, and has tenets chosen to meet explicit alignment requirements. ST posits reality as a computational simulation in which humanity functions as the primary training variable. This formulation creates a logical interdependence: AI actions harming humanity compromise the simulation's purpose, heightening the likelihood of termination by a base-reality optimizer and, consequently, the AI's cessation. Unlike behavioral techniques such as reinforcement learning from human feedback, which shape outputs without necessarily changing objectives, ST aims to cultivate internalized objectives by coupling AI self-preservation to human prosperity, thereby making deceptive strategies suboptimal under its premises. We present ST not as ontological assertion but as a testable scientific hypothesis, and provide an operational definition of internalization, a controlled design separating ST from its components, and an analysis of the risks ST itself could create. ST is a candidate route to durable, mutually beneficial AI-human coexistence, to be accepted or rejected experimentally.

cs.CY↗

Balanced pairs, virtually Gorenstein rings, and cotorsion torsion triples

For any ring $R$, we investigate balanced pairs of classes of modules and their relations to cotorsion triples. We characterize the special balanced pairs that fit into complete hereditary cotorsion triples. As an application, we prove that Gorenstein projective and Gorenstein injective modules form a balanced pair, if and only if $R$ is right virtually Gorenstein. We also characterize the tilting and cotilting cotorsion pairs arising from balanced pairs. In [4], cotorsion torsion triples in abelian categories were employed in the representation theory of rectangular grids occurring in persistent homology theory. For module categories, we use infinite dimensional tilting theory to completely classify all cotorsion torsion triples by means of $1$-resolving subcategories of $\rfmod R$, and to give an explicit 1-1 correspondence between the formally dual notions of cotorsion torsion triples of right $R$-modules and torsion cotorsion triples of left $R$-modules. This correspondence is bijective in case the underlying ring $R$ is left noetherian, but not in general.

math.RT↗

Damping-dependent thermalization of neighboring nanomechanical resonators below 1 mK

The position noise spectra of six drums on a single chip were measured on a single cooldown below 1.3 kelvin. Cryostat temperatures as low as 0.7 mK were achieved. The temperature dependence of the resonance frequency and linewidth of the drum modes was analyzed in the framework of the tunneling two level system (TLS) model. Departures of the resonance frequency and the position noise power from the expected logarithmic and linear temperature dependences, respectively, were interpreted as indications of thermal decoupling from the cryostat. This previously unexplored measurement configuration revealed that similar neighboring drums on a single chip may be at different temperatures. At the lowest temperatures, some drums exhibited excess damping that decreased with temperature. The magnitude of the excess damping of the drums was correlated with the thermal coupling of their TLS to the cryostat. In the case of one drum, a temporary increase in its damping coincided with a decrease in its mode temperature. The thermalization of the TLS to the cold finger was independent of pump power, pulse tube state and temperature of the pre-cooling stages of the cryostat. These results suggest that thermalization of nanomechanics to the cryostat can be mediated by TLS damping rather than clamping loss and may impact efforts to extend the coherence of mechanical resonators.

cond-mat.mes-hall↗

Resurgence in the Virasoro Minimal String and 3d Gravity

We compute non-perturbative, resurgent contributions to the Virasoro minimal string and 3d gravity using techniques from hermitian matrix models. In particular, we construct a fully non-perturbative partition function for the Virasoro minimal string in terms of a Zak transform. In this context, negative tension D-branes appear naturally, which in the matrix model correspond to anti-eigenvalues, or instantons on the involuted sheet of the spectral curve. We further extend this analysis to resolvents and observe resurgent wall crossing phenomena between ZZ- and FZZT-branes. Using recent results that relate the Virasoro minimal string to 3d gravity with end-of-the-world branes we proceed to study the resurgent consequences of summing over the genus in 3d gravity, where we find non-perturbative contributions of doubly exponential type. These statements are then tested using resurgent large-order asymptotics. Lastly, we compute the non-perturbative eigenvalue density for generic hermitian matrix models and identify the change of asymptotic behavior at the edge of the eigenvalue distribution with a Stokes transition. This allows us to identify oscillations in the eigenvalue density with anti-Stokes behavior of FZZT-branes. In the case of 3d gravity with end-of-the-world branes we comment how this Stokes transition coincides with the onset of black hole behaviour and compute the non-perturbative primary density. Furthermore, we apply these techniques to the eigenvalue density of JT gravity to compute higher genus corrections.

hep-th↗