Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 271 records · Page 15Linked to original sources

GenGait: A Transformer-Based Model for Human Gait Anomaly Detection and Normative Twin Generation

Gait analysis provides an objective characterization of locomotor function and is widely used to support diagnosis and rehabilitation monitoring across neurological and orthopedic disorders. Deep learning has been increasingly applied to this domain, yet most approaches rely on supervised classifiers trained on disease-labeled data, limiting generalization to heterogeneous pathological presentations. The methodological objective of this work is to develop a label-free framework for joint-level anomaly detection and kinematic correction based on a Transformer masked autoencoder trained exclusively on normative gait sequences from 150 adults, acquired with a markerless multi-camera motion-capture system. At inference, a two-pass procedure is applied to potentially pathological input sequences: first, it estimates joint inconsistency scores by occluding individual joints and measuring deviations from the learned normative prior. Then, it withholds the flagged joints from the encoder input and reconstructs the full skeleton from the remaining spatiotemporal context, yielding corrected kinematic trajectories at the flagged positions. The validation objective is to assess whether the framework preserves unseen normative gait and reduces angular deviation in simulated abnormal gait patterns. In this proof-of-concept evaluation, data from 10 held-out normative participants, who performed seven simulated abnormal gait patterns, showed a significant reduction in angular deviation across all analyzed joints with large effect sizes, and preservation of normative kinematics. The proposed approach enables interpretable, subject-specific localization of joints that are inconsistent with learned normative gait patterns and generation of an individualized normative reconstruction without requiring disease labels. Video is available at https://youtu.be/Rcm3jqR5pN4.

cs.AI↗

High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination

Humans exhibit remarkable abilities to coordinate in groups. As large language models (LLMs) become more capable, it remains an open question whether they can demonstrate comparable adaptive coordination and whether they use the same strategies as humans. To better understand this, we compare LLM and human performance on a common-interest game with imperfect monitoring: Group Binary Search. In this $n$-player game, participants need to coordinate their actions to achieve a common objective. Players independently submit numerical values in an effort to collectively sum to a randomly assigned target number. Without direct communication, they rely on group feedback to iteratively adjust their submissions until they reach the target number. Our findings show that, unlike humans who adapt and stabilize their behavior over time, LLMs often fail to improve across games and exhibit excessive switching, which impairs group convergence. Moreover, richer feedback (e.g., numerical error magnitude) benefits humans substantially but has small effects on LLMs. Finally, we show that GRPO can be effective in reducing the excessive switching. Taken together, by grounding the analysis in human baselines and mechanism-level metrics, including reactivity scaling, switching dynamics, and learning across games, we point to differences in human and LLM groups and provide a behaviorally grounded diagnostic for closing the coordination gap.

cs.MA↗

Understanding Latent Diffusability via Fisher Geometry

Diffusion models often degrade in latent spaces, yet the formal causes remain poorly understood. We quantify latent-space diffusability via the rate of change of the Minimum Mean Squared Error (MMSE) along the diffusion trajectory. Our framework decomposes this MMSE rate into contributions from Fisher Information (FI) and Fisher Information Rate (FIR). We show that isometric embeddings preserve intrinsic FI and establish quantitative intrinsic-FI bounds for a broader class of bi-Lipschitz encoders with controlled weak volume distortion, whereas FIR is governed by the interplay between encoder and data geometries. Our analysis separates four geometric contributions in local stability bounds for Gaussian-smoothed FIR: dimensional compression, tangential distortion, high-frequency encoder curvature, and curvature of data manifold. Experiments across diverse autoencoding architectures provide qualitative support for the geometric mechanisms identified by the theory and show that empirical FI and FIR track several measures of generation quality and latent-space geometry in the settings tested. We establish FI and FIR as a comprehensive analytical framework for understanding latent diffusability.

cs.LG↗

On the instability of some upward propagating, exact, nonlinear mountain waves

Using the short-wavelength instability method, we investigate the linear instability of an exact solution describing upward-propagating mountain waves, derived in A.~Constantin, \emph{J. Phys. A: Math. Theor.} (2023), for a dry and adiabatic compressible flow. Within this approach, the stability problem reduces to the analysis of a system of ordinary differential equations with periodic coefficients along the fluid trajectories, which we study by means of Floquet theory. We show that the Floquet multipliers are of the form $\{1,μ,μ^{-1}\}$, the multiplier $1$ being associated with the conservation of potential vorticity, so that stability is decided by a single discriminant, which coincides with that of a Hill equation. For neutral stratification the flow is unstable whenever the wave steepness exceeds the critical value $\frac{1}{3}$, while a weakly stable stratification lowers this threshold to $\frac{1}{3}-\frac{4}{9}ν^2+O(ν^4)$, $ν$ being the ratio between the Brunt--Väisälä frequency and the orbital frequency of the air parcels. The instability is a subharmonic parametric resonance. Since the solution is given in Lagrangian coordinates, these results identify an unstable region in the upper part of the laminar layer, beneath the tropopause, where the wave is steepest.

physics.ao-ph↗

Two-stage disruption of resonant chains

TESS is enabling the discovery of transiting planets around young stars. These observations suggest that most close-in planets were born in chains of mean-motion resonances that break on a characteristic timescale of order 100 Myr. This observation is surprising because the same dissipative forces that capture planets into resonance render their orbits long-term stable. We explore a two-stage disruption scenario for resonant chains of super-Earths. First, the chains have their (free) eccentricities excited by some mechanism. We show that any such mechanism that seeds eccentricities of a few percent sets in motion a second stage of dynamical instability on a $\sim$100 Myr timescale. A possible stage-one mechanism is the accretion of a handful of Mercury-sized bodies totaling a few percent of the planetary system mass, which excites the requisite eccentricities and triggers a stage two that reproduces the observed decline in the incidence of resonance. Impacts from such bodies can also explain why some young systems have period ratios narrow of commensurability. We sketch how these impactors may have grown out of debris left over from an earlier epoch of planet formation. We also identify two new trends in the observational data: a decline in multiplicity on the same timescale as the decline in the incidence of resonance, and an increase in the occupation of resonances with multiplicity.

astro-ph.EP↗

Optimal Centered Active Excitation in Linear System Identification

We propose an active learning algorithm for linear system identification with optimal centered noise excitation. Notably, our algorithm, based on ordinary least squares and semidefinite programming, attains the minimal sample complexity while allowing for efficient computation of an estimate of a system matrix. More specifically, we first establish lower bounds of the sample complexity for any active learning algorithm to attain the prescribed accuracy and confidence levels. Next, we derive a sample complexity upper bound of the proposed algorithm, which matches the lower bound for any algorithm up to universal factors. Our tight bounds are easy to interpret and explicitly show their dependence on the system parameters such as the state dimension.

math.OC↗

Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR

Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR architectures connect a speech encoder to a large language model via a projection module, enabling adaptation with text-only data. However, this introduces a modality gap, as the LLM is not exposed to the noisy representations produced by the speech projector. We investigate whether small amounts of speech can mitigate this mismatch. We compare three strategies: text-only adaptation, paired speech-text adaptation, and mixed batching (MB), which combines both. Experiments in in-domain and out-of-domain settings show that even limited speech consistently improves performance. Notably, MB using only 10% of the target-domain (less than 4 hours) speech achieves word error rates comparable to, or better than, conventional ASR fine-tuning with the full dataset, indicating that small amounts of speech provide a strong modality-alignment signal.

cs.CL↗

What Does a Sharing Question Add? Auditing LLM Survey Scores for Misinformation

Evaluating misinformation requires distinguishing whether readers believe content from whether they would share it. Asking large language models (LLMs) both questions yields two scores, but does the sharing answer contribute information beyond the credibility answer? We audit eight model versions on 290 synthetic misinformation articles, using 1,256 paired survey responses with 317 participant identifiers as an external validity criterion. An initial reversal motivates the audit: every model's raw sharing score predicts mean human sharing less accurately than its credibility score. This ordering changes after offset correction, so it does not by itself diagnose missing information. We instead distinguish score reconstructability, persistence across elicitation formats, and incremental human validity. Credibility predicts 30.8-72.5% of model-sharing variation relative to a held-out constant baseline; remaining sharing differences correlate at 0.61-0.75 across question-order and separate-question conditions. Yet adding model sharing to human and model credibility yields only -0.25% to +0.69% error reduction with fixed regression, with all exploratory intervals crossing zero. Flexible prediction and format changes do not establish an improvement. Sharing answers therefore contain structured variation beyond the observed credibility score, without established incremental validity for human sharing in these data. The findings motivate validating the contribution of each elicited outcome, beyond inspecting score differences or agreement across prompts.

cs.AI↗

SAFE: Spatially-Aware Feedback Enhancement for Fault-Tolerant Trust Management in Event-Based VANETs

In event-based trust management for vehicular ad hoc networks (VANETs), vehicles that witness a road event broadcast its state, and vehicles that later witness the same event score the earlier broadcasters in feedback reports sent to a central decision unit (CDU). When the event state changes, an honest vehicle that reported the previous state receives negative feedback, as the evaluator compares an out-of-date message with the new state. This stale-witness problem is hypothesised to be a major source of unfair penalisation of honest vehicles. To address it, SAFE (Spatially-Aware Feedback Enhancement) is proposed. In SAFE, vehicles continue to record event messages after their decision and throughout the witness area, and send an updated feedback report when they leave it. SAFE was compared with the trust cascading-based emergency message dissemination model (TCEMD) in attack-free highway scenarios simulated with OMNeT++, Veins and Simulation of Urban MObility (SUMO). In the single-event scenario, the negative-feedback rate in the two update rounds after the state change decreased from 53.8% to 24.7% and from 55.6% to 8.3%, and the number of distinct honest vehicles blacklisted decreased from 20 to 6. In the multi-event scenario, the negative-feedback rate remained at or below 1.3% in SAFE, compared with up to 77.6% in TCEMD, and the share of trust evaluations ending in an untrusted label decreased from 22.7% to at most 1.7%. These gains required 2.3 to 5.0 times more feedback entries. Experiments with two decision distances showed that a shorter gap between the decision and witness distances reduced stale feedback in both schemes, supporting the stale-witness hypothesis

cs.NI↗

Exponential quantum advantage in processing massive classical data

Broadly applicable quantum advantage, particularly in classical data processing and machine learning, has been a fundamental open problem. In this work, we prove that a small quantum computer of polylogarithmic size can perform large-scale classification and dimension reduction on massive classical data by processing samples on the fly, whereas any classical machine achieving the same prediction performance requires exponentially larger size. Furthermore, classical machines that are exponentially larger yet below the required size need superpolynomially more samples and time. We provide evidence for these quantum advantages in real-world applications, including single-cell RNA sequencing and movie review sentiment analysis, demonstrating four to six orders of magnitude reduction in size with fewer than 60 logical qubits. These quantum advantages are enabled by quantum oracle sketching, an algorithm for accessing the classical world in quantum superposition using only random classical data samples. Combined with classical shadows, our algorithm circumvents the data loading and readout bottleneck to construct succinct classical models from massive classical data, a task provably impossible for any classical machine that is not exponentially larger than the quantum machine. These quantum advantages persist even when classical machines are granted unlimited time or if BPP = BQP, and rely only on the correctness of quantum mechanics. Together, our results establish machine learning on classical data as a broad and natural domain of quantum advantage and a fundamental test of quantum mechanics at the complexity frontier.

quant-ph↗

Stochastic Thermodynamics for Autoregressive Generative Models: A Non-Markovian Perspective

Autoregressive generative models -- including Transformers, recurrent neural networks, classical Kalman filters, state space models, and Mamba -- all generate sequences by sampling each output from a deterministic summary of the past, producing genuinely non-Markovian observed processes. We develop a general theoretical framework based on stochastic thermodynamics for this class of architectures and introduce the entropy production, which can be efficiently estimated from sampled trajectories without exponential overhead, despite the non-Markovian nature of the observed dynamics. As a proof-of-concept experiment with a large language model (LLM), we evaluate the entropy production for a pre-trained Transformer-based model, GPT-2. We find that the token-level entropy production is dominated by a syntactic artifact, while the sentence-level entropy production tends to be larger for causally ordered than for non-causal text sets. This observation is supported by a re-evaluation with a substantially larger model, Qwen3-4B-Base. We also demonstrate the framework in the linear Gaussian case, where the model reduces to the Kalman innovation representation and the entropy production admits an analytical expression. We also show that the entropy production decomposes exactly into non-negative per-step contributions in terms of retrospective inference, and each of those terms further splits into information-theoretically meaningful terms: a compression loss and a model mismatch. Our results establish a bridge between stochastic thermodynamics and modern generative models, and provide a starting point for using irreversibility as a quantitative probe of the highly non-Markovian processes generated by models such as LLMs.

cond-mat.stat-mech↗

Cohesion from Action-Dependent Perception in Decision-Based Active Matter

Many biological and artificial agents reorient by locally optimizing a utility over a finite set of candidate actions. We show that decision rules based on pairwise additive scores over candidate-dependent interaction neighborhoods provide a generic mechanism for a cohesive bias in active matter: the utility factorizes into the product of an average interaction score and the number of recruited neighbors. Maximizing it therefore couples the decision to neighbor recruitment as a structural consequence of the decision process itself. For alignment-based utilities and orientational candidate actions, our models reduce to a variant of the Vicsek model when the interaction neighborhood is independent of the candidate action. This identifies the mechanism behind our previous finding that maximizing predicted future alignment produces cohesive flocks, and extends it to arbitrary candidate sets and additive utilities, provided that at least some actions can have positive scores. We further study scanning alignment, in which agents evaluate a topological neighborhood within a finite vision cone along each candidate heading, without future-state predictions. Even the baseline variant---with a sign-changing score and zero-neighbor candidates eligible---already produces nontrivial cohesive flocks and swarms in open, unbounded two-dimensional space. Excluding zero-neighbor candidates, or replacing the score with a non-negative one, yields substantially richer collective behavior that additionally includes swirling. We focus on the former as the most biologically motivated choice. The swirling states occupy an interval of noise intensities and spontaneously reverse their sense of rotation; the rotation appears to set in abruptly at the low-noise edge and to die away gradually at the high-noise one.

cond-mat.soft↗

Superconducting orbital diode effect in SN bilayers

We study the superconducting diode effect (SDE) in a diffusive superconductor - normal metal (SN) bilayer subjected to an in-plane magnetic field. The supercurrent flows along the layers, perpendicular to the field. The SDE, manifested as an asymmetry in the critical (depairing) currents and kinetic inductance for opposite current directions, arises from an orbital mechanism due to the inhomogeneous distribution of the Meissner currents caused by a spatially varying superfluid density. Recently, Levichev et al. [Phys. Rev. B 108, 094517 (2023)] demonstrated the realization of this effect in such a structure, supporting numerical calculations for an ideal interface with an experiment. In this work, we investigate the influence of a nonideal interface with finite resistance on the SDE. Employing an analytical approach, we focus on limiting cases corresponding to weak intralayer inhomogeneities. We find that the strength of the SDE depends nonmonotonically on the interface resistance when the bilayer thickness is small compared to the coherence length. Remarkably, a nonideal interface can enhance the SDE compared to the ideal case. We also demonstrate that, in the vicinity of the phase transition, a regime of unidirectional superconductivity can be realized, in which superconductivity exists only within an interval of finite currents flowing in the same direction.

cond-mat.supr-con↗

Numerically optimized amplitude-robust controlled-Z gate for ultracold neutral atoms with individual addressing capability

We numerically optimized a scheme for a symmetric controlled-Z (CZ) gate based on the Rydberg blockade in neutral atoms, with the aim of increasing its robustness to variations in the Rabi frequency. The scheme exploits analytically defined phase profiles of the laser pulse and proves almost an order of magnitude more robust to Rabi-frequency variations than previously proposed protocols. We designed this amplitude-robust protocol with a smooth pulse shape and optimized it for a finite Rydberg blockade strength, which makes it suitable for experimental implementation in both single-photon and two-photon configurations. We further show that the protocol can be applied to individually addressed Rydberg excitation, accounting for the asymmetry of the Rabi frequencies of two atoms excited by tightly focused laser beams. This allows the effects of residual thermal motion of the trapped atoms to be reduced. Finally, we examined the performance of the protocol for single-photon and two-photon Rydberg excitation schemes at finite blockade strength, and demonstrated its advantages for individual addressing at finite temperatures of the trapped atoms.

quant-ph↗

Hierarchical Secure Distributed Linearly Separable Computation with Arbitrary Heterogeneous Data Assignment

This paper studies secure distributed linearly separable computation over a three-layer hierarchical network, where clustered users communicate with a central server through relays. The server aims to recover Kc linear combinations of K intermediate outcomes, where each intermediate outcome is a separable function of one dataset. We consider a more general setting with arbitrary heterogeneous data assignment across users, where `arbitrary' means that the data assignment is given in advance (which can be in any form) and `heterogeneous' means that the users may hold different numbers of datasets. Under this assignment, each user computes the intermediate outcomes of its assigned datasets and sends masked messages to its associated relay. The relays subsequently process and forward the received messages to the server. We impose two security constraints: (i) security against server, requiring the server to learn only the desired task function without gaining any additional information about users' inputs; and (ii) security against relays, ensuring each relay learns nothing about users' inputs. Moreover, the server or any relay may collude with a subset of users. For Kc=1, the underlying computation reduces to distributed gradient coding. We propose a secure scheme tolerating user dropouts and user collusion, achieving the optimal two-layer communication rates in one regime and order-optimal communication rates within a factor of 2 in the other regime. For Kc>1, we extend the proposed construction to multi-dimensional linearly separable tasks under the no-dropout setting.

cs.IT↗

From Order to Distribution: An Exact Operator Framework for Forgetting in Continual Learning

A central challenge in continual learning is forgetting: the loss of performance on previously learned tasks after learning new ones. Prior theory has analyzed forgetting under random orderings of fixed task collections in overparameterized linear regression. We shift the focus from task order to task distribution, asking how its structure determines forgetting. In the linear setting with a shared solution, i.i.d. task sampling, and sequential exact fitting, we derive an exact operator identity expressing historical forgetting directly in terms of the task distribution. Building on this identity, we establish an exponential decay guarantee for expected historical forgetting under every fixed task distribution in finite dimensions, characterize its asymptotic behavior, and relate decay to the distribution's coverage of observable directions. For an individual learned task, we show that subsequent tasks can collectively support recovery without exact revisits. We derive a lower bound on recovery time and construct a task distribution attaining its inverse-coverage scaling.

cs.LG↗

Daycare Matching with Siblings: Social Implementation and Welfare Evaluation

In centralized matching markets, agents may value joint assignment, as with siblings or couples. Standard preference estimation ignores such complementarities, complicating welfare analysis of priority rules for paired assignment. We develop an empirical framework incorporating these preferences and apply it to Japanese daycare assignment. Families face both additional commuting distance and a fixed disutility from split assignment. We estimate the latter at 4.61 commuting-kilometer equivalents. Our fixed-report counterfactual estimates that the reform increased mean welfare by 0.032 kilometer-equivalent units. Ignoring sibling complementarity understates welfare gains for households applying simultaneously for multiple children by about 27%.

econ.GN↗

Multi-User mmWave Beam and Rate Adaptation via Combinatorial Satisficing Bandits

We study downlink beam and rate adaptation in a multi-user mmWave MISO system where multiple base stations (BSs), each using analog beamforming from finite codebooks, serve multiple single-antenna user equipments (UEs) with a unique beam per UE and discrete data transmission rates. BSs learn about transmission success based on ACK/NACK feedback. To encode service goals, we introduce a satisficing throughput threshold $τ_r$ and cast joint beam and rate adaptation as a combinatorial semi-bandit over beam-rate tuples. Within this framework, we propose SAT-CTS, a lightweight, threshold-aware policy that blends conservative confidence estimates with posterior sampling, steering learning toward meeting $τ_r$ rather than merely maximizing. Our main theoretical contribution provides the first finite-time regret bounds for combinatorial semi-bandits with satisficing objective: when $τ_r$ is realizable, we upper bound the cumulative satisficing regret to the target with a time-independent constant, and when $τ_r$ is non-realizable, we show that SAT-CTS incurs only a finite expected transient outside committed CTS rounds, after which its regret is governed by the sum of the regret contributions of restarted CTS rounds, yielding an $O((\log T)^2)$ standard regret bound. On the practical side, we evaluate the performance via cumulative satisficing regret to $τ_r$ alongside standard regret and fairness. Experiments with time-varying sparse multipath channels show that SAT-CTS consistently reduces satisficing regret and maintains competitive standard regret, while achieving favorable average throughput and fairness across users, indicating that feedback-efficient learning can equitably allocate beams and rates to meet QoS targets without channel state knowledge.

cs.LG↗