Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI

Human label variation in natural language inference is increasingly treated as signal rather than noise, but how much of it formal semantic structure explains has not been measured directly. We measure it on the 3,113 SNLI and MNLI items of ChaosNLI, using a rule-based operator and monotonicity tagger validated against MED (0.883 agreement at the edit site, 0.807 on the sentence-level summary our analyses consume), three preregistered analysis blocks, and full reporting of negative results. Three bounds emerge. First, a group-level boundary: hypotheses that are not purely upward monotone show reliably higher label entropy (Cliff's delta = -0.284), and rank-based tests defend the effect against operator-presence and length reductions, though a bounded-outcome sensitivity check weakens the regression form of the length defense. Second, an item-level ceiling: the same formal profiles explain only 3.3 to 3.6 percent of entropy variance and reach a median-split AUC of 0.606, too weak to identify high-disagreement items. Third, composition invariance: across the boundary, three high-powered preregistered contrasts on validated error shares and explanation-type shares (VariErr, LiTEx) all return null results. In this sample, formal semantic structure shifts how much annotators disagree by a small amount and does not detectably change what they disagree about. ChaosNLI-S/M consists of items selected for low original agreement, and every claim is conditioned on that scope. All analyses were preregistered in a version-controlled research log, whose audit trail, including one corrected interpretation rule, the paper discloses.

cs.CL↗

Portal Matter and Scotogenic-like Dirac Neutrino Masses

Loops of portal matter (PM) fields, carrying both Standard Model (SM) and dark charges, can generate the necessary kinetic mixing (KM) between the ordinary and dark photons (DP) in vector portal scenarios thus allowing for interactions between visible and dark sector fields. Here we show that the field content of a previously considered model based on a partial $E_6$-like UV-completion of such setups can generate light Dirac neutrino masses in the interesting range, $\sim 0.05$ eV, at the one-loop level similar to what happens in scotogenic dark matter (DM) scenarios. While general $E_6-$like PM scenarios have been shown to be easily probed at colliders, such as the HL-LHC, uniquely testing this specific subclass of these setups in a direct fashion is found to be somewhat more challenging.

hep-ph↗

RegionFM: Interpretable Region-Based Brain MRI Classification Using Foundation Model Embeddings

Foundation models provide powerful representations for brain MRI analysis, but their predictions remain difficult to interpret in anatomically meaningful terms. Clinical assessment of brain MRI is commonly organized around anatomically defined structures and regional abnormalities, whereas conventional explanation methods typically produce voxel- or patch-level importance maps that do not explicitly quantify the contributions of individual brain regions. To address this mismatch, we propose RegionFM, an interpretable framework that integrates anatomical segmentation with brain MRI foundation-model embeddings. RegionFM first divides each MRI scan into anatomical regions and constructs a separate MRI volume for each region. A frozen foundation model then encodes each region into an embedding, and a region-additive logistic model combines these embeddings such that every anatomical region contributes an explicit scalar term to the final prediction. This formulation supports both subject-level and cohort-level analyses of regional contributions. We evaluate RegionFM on cognitive-impairment classification using embeddings from multiple pretrained brain MRI foundation models. The results show that RegionFM maintains performance comparable to less interpretable fine-tuning approaches while providing anatomically grounded explanations. Randomized embedding ablations yield near-chance performance, indicating that the predictions rely on meaningful structure captured by the foundation-model embeddings rather than simple feature statistics. Overall, RegionFM better aligns model explanations with anatomy-based clinical reasoning while maintaining competitive predictive performance.

cs.CV↗

An affine Birkhoff-Kellogg type result in product spaces and its application to differential systems

We prove a version of the Birkhoff-Kellogg theorem in product spaces, under affine transformations. This is a fairly natural framework that arises when dealing with parameter-dependent systems of functional-differential equations. We provide criteria on the existence of parameter-dependent solutions that have all the components nontrivial. The theoretical results are illustrated in the case of systems of second order functional-differential equations, where the functional part can cover the interesting cases of time and state-dependent deviated arguments. We provide a concrete example, for which we furnish numerically approximated solutions, coherent with our theoretical framework, and in which we compute or estimate all the constants that are required by our abstract results.

math.DS↗

Tensor-Based Dynamic Channel Estimation for mmWave Movable Antenna MIMO Systems

This paper investigates the dynamic channel estimation algorithm in mmWave movable antenna (MA) multiple-input multiple-output (MIMO) systems. To achieve highly accurate channel estimation, we propose a tensor decomposition-based channel estimation algorithm. First, by leveraging the path response model and utilizing the intrinsic sparsity of mmWave channels, the channel corresponding to MA pairs at the base station and mobile station is transformed into a superposition of channels from sparse paths. Next, the received signal is constructed as a fourth-order tensor to fully capture the high-dimensional structural information of the MA MIMO channel. Then, two tensor decomposition schemes are adopted to extract the factor matrices, and our analysis reveals that the uniqueness of the decomposition can be guaranteed in our model. Subsequently, the propagation loss, frequency offset, angle of arrival/departure, and time delay are obtained based on these factor matrices and the channel matrix can be rebuilt. Additionally, Cramér-Rao bound (CRB) is also derived as a performance evaluation standard, proving that the proposed algorithm achieves a higher estimation accuracy and nearly approaches this minimum bound. Moreover, normalized mean square error (NMSE) is selected as the evaluation metrics for estimation accuracy. Finally, simulation results reveal a notable reduction in the estimation error of the proposed algorithm when compared to the baseline algorithms, confirming its estimation advantage.

eess.SP↗

Equilibrium Play Without Mutual Knowledge of Rationality

Equilibrium play in two-player zero-sum games is usually justified via epistemic assumptions, such as mutual knowledge of rationality and beliefs, that go far beyond the rationality of the players. We propose a justification that dispenses with these assumptions. To this end, we consider solution concepts that assign to every subgame of a given game a set of plausible actions for each player, and we impose two conditions. Rationality requires that the plausible sets are supports of undominated strategies or, equivalently, that all plausible actions are best responses to a common belief about the opponent. Inheritance requires that plausible actions remain plausible when implausible actions are discarded. In two-player games, the two conditions generically characterize the solution concepts that consistently select supports of Nash equilibria. Zero-sum payoffs ensure that Nash equilibria are generically unique, so that such a selection exists and is unique. Equilibrium play thus emerges from individual rationality and the mutual understanding that plausibility judgments persist when implausible actions are discarded.

econ.TH↗

On the structure of $2$-step nilpotent Lorentzian naturally reductive Lie groups

We study $2$-step nilpotent Lie groups with naturally reductive left-invariant Lorentzian metrics with respect to the presentation group $N\rtimes H^{\operatorname{aut}}$. Replacing the standard non-degenerate center assumption with the weaker condition that the commutator ideal be non-degenerate, we develop a framework that extends the representation-theoretic construction to the Lorentzian setting and covers both the non-degenerate and degenerate center cases. In the degenerate case, we show that the associated Lie algebra is a central extension of a semidirect product whose Riemannian factor is naturally reductive. Furthermore, we obtain invariant decompositions of the defining representation in the non-degenerate case. In the degenerate case, we introduce a Riemannian quotient representation and show that the kernel of this representation is a central abelian ideal whose elements act by null rotations. Finally, we use these results to characterize the isotropy algebra of the group of isometric automorphisms.

math.DG↗

Using Fine-Tuned LLMs to Identify Indicators of Vulnerability in UK Police Incident Logs

Purpose: Understanding how much of routine policing involves vulnerable people could inform resourcing, training, and multi-agency response, yet administrative data provide limited insight. We explore whether an LLM-based classification pipeline, developed on open-source US police data, can be adapted to estimate the prevalence of four vulnerability indicators - mental ill health, substance misuse, alcohol dependence, and homelessness - in UK police incident narratives, and when outputs can be treated as defensible measurements. Methods: We analyse nearly 3,000 de-identified incident logs from a UK police force, using a multi-stage pipeline combining repeated model inference, label aggregation, structured human review, and statistical correction. The pipeline runs on a locally hosted open-weight LLM, reflecting the secure environments police must work in. Results: LLMs can produce meaningful, if imperfect, prevalence estimates at scale. Mental ill health indicators are present in approximately one in five incidents, with lower prevalence for other indicators. However, naive LLM deployment is unreliable: single-pass classifications are unstable, and aggregated outputs systematically over-assign indicators relative to human judgement. Correcting these biases required substantial human input and statistical adjustment, leaving considerable uncertainty. Conclusions: While LLMs can extract information from unstructured police data, their outputs cannot be treated as valid measurements without careful methodological support. At the population level, defensible estimates are achievable but resource-intensive; at the individual level, errors remain frequent and unpredictable, limiting suitability for operational decisions. This study highlights both the potential and the constraints of LLM-based measurement in applied settings.

cs.CL↗

Contrastive On-Policy Distillation

On-policy distillation (OPD) trains a student model on trajectories sampled from its own policy, providing dense token-level supervision by minimizing the divergence between the teacher's and student's output distributions at each token. Although existing OPD approaches effectively distill strong reasoning capabilities into student models, they inherently inherit uncurated chain-of-thought traces, exacerbating overthinking and reasoning redundancy. To address this limitation, we introduce COPD, a contrastive OPD framework. Specifically, for each token generated by the student, a frozen teacher evaluates the current student state under two contrasting prompts that induce low and high reasoning effort, respectively. The resulting difference in log-probabilities serves as a token-level advantage signal to guide the policy update. Rather than merely mimicking a single teacher distribution, COPD guides the student model toward acquiring concise and efficient reasoning strategies. We evaluate COPD across 9 multimodal benchmarks covering both reasoning and understanding tasks. Empirical results show that COPD substantially reduces reasoning length without hurting task performance, consistently improving efficiency across different tasks and model scales. In addition, this contrastive paradigm extends naturally to on-policy self-distillation (OPSD), establishing self-contrastive supervisory signals that enable a single model to compress its own reasoning without an external teacher.

cs.CV↗

Enhancing Relation Modeling with Social Attributes for Social Media Popularity Prediction

Recent studies highlight the critical role of retrieval-augmented mechanisms in social media popularity prediction (SMPP). Although such frameworks have improved SMPP performance by leveraging historical posts, existing methods still suffer from the low retrieval accuracy due to the oversight of relative relationships among UGC instances. To address this limitation, we propose a novel Relation-Enhanced Retrieval-Augmented framework (RE-Rag) that models UGC similarity as a continuous relation jointly driven by semantic content and social attributes. Specifically, RE-Rag employs a Semantic-Attribute Retriever (SAR) to obtain instances aligned in both semantic and social-attribute distributions. Subsequently, we design a Relation-Guided Predictor (RGP): first, cross-attention encodes multimodal features of retrieved instances; then, a relative relation graph is introduced to guide attention weight allocation, forming a Relation-Guided Transformer (RGTs) that dynamically modulate attention weights based on relative attribute relations to capture the interplay between semantics and various social attributes. The refined features are fused with the target instance for popularity prediction. Experiments on three public benchmarks show that RE-Rag consistently outperforms state-of-the-art methods in both prediction accuracy and retrieval efficiency.

cs.MM↗

Schrödinger perturbation theory for black hole quasinormal modes

Deviations from vacuum general relativity (such as modified theories or the presence of an environment) produce small shifts in black hole quasinormal mode (QNM) spectra. These effects are becoming increasingly relevant for gravitational wave astronomy as observations of ringdown spectra become more precise. To leading order in deviations from vacuum general relativity, frequency shifts can be calculated using eigenvalue perturbation theory. However, for Kerr, there is no systematic framework to compute higher order corrections. The major obstacle is that QNMs do not form a complete basis due to the non-self-adjointness of the system. Nevertheless, it was recently shown that QNMs are orthogonal with respect to an appropriate bilinear form. In this work, we use the bilinear form to systematically lift Schrödinger perturbation theory to the black hole setting. We obtain a formula for quasinormal frequency shifts to any order, in terms of lower order mode shifts. We also provide a spectral decomposition of the first-order mode shift, which involves projections onto unperturbed QNMs along with continuous-spectrum contributions---making incompleteness explicit. We illustrate the framework on slowly-spinning Kerr and Pöschl-Teller examples, where we find that the QNM sum itself diverges.

gr-qc↗

Latent-Action-Guided Vision-Language Contrastive Learning for Surgical Interaction Recognition

Recognizing instrument-tissue interactions is essential for context-aware surgical AI. Vision-language models offer a natural way to inject semantic structure into surgical representations by aligning video features with textual action descriptions. However, pretrained encoders may lack spatial coherence, while global semantic alignment does not ensure precise spatial and temporal representations. By analyzing frame-to-frame feature changes, we find that semantic alignment increases their dimensionality, but larger increases do not necessarily improve recognition; encoders also differ in how strongly dominant changes localize to interaction regions. Motivated by these findings, we introduce LAViFiT, which compresses frame-to-frame changes into latent actions and predicts next-frame features during end-to-end video-language alignment. Without additional spatial or motion annotations, LAViFiT improves the interaction grounding of leading feature changes and temporal-direction sensitivity in our evaluated settings. We further characterize how action capacity and prediction strength affect recognition across encoders and triplet components. Using image encoders without large-scale video pretraining, LAViFiT achieves competitive recognition with faster inference and smaller INT4 accuracy drops than V-JEPA2/2.1, supporting its deployment potential.

cs.CV↗

Skeleton Chordalities

We study new higher-dimensional analogs of graph chordality and review the existing ones. Our main results for simplicial complexes are: (1) $Δ$ skeleton-E-chordal $\Rightarrow$ $Δ^\vee$ vertex-decomposable $\Rightarrow$ $Δ$ skeleton-clique-chordal. Moreover, for subflag complexes, $Δ$ skeleton-E-chordal $\Longleftrightarrow$ $Δ^\vee$ vertex-decomposable. (For $d=1$ this boils down to ``$G$ chordal $\Longleftrightarrow$ $G^\vee$ vertex-decomposable'', a result closely related to Fröberg's theorem.) (2) For subflag complexes, $Δ$ is skeleton-E-chordal $\Longleftrightarrow$ it splits as $Δ= Δ_1 \cup Δ_2$, with each $Δ_i$ a skeleton-E-chordal induced subcomplex of $Δ$, and with $Δ_1 \cap Δ_2$ a complex whose $1$-skeleton is a clique. (This generalizes ``$G$ chordal $\Longleftrightarrow$ $G$ splits as a union of chordal graphs that intersect in a common clique''). (3) $Δ$ skeleton-E-chordal $\Longleftrightarrow$ every nonempty induced subcomplex of $Δ$ has a skeleton-E-simplicial vertex. (Generalizes ``$G$ chordal $\Leftrightarrow$ every nonempty induced subgraph has a simplicial vertex''.) (4) $Δ$ underclosed $\Rightarrow$ $Δ$ skeleton-weakly-chordal and weakly-closed. (Generalizes ``$G$ interval $\Rightarrow$ $G$ chordal and co-comparability''.) (5) All pure E-chordal complexes are vertex-chordal; all pure mid-chordal complexes are weakly-vertex-chordal; all pure very-weakly-chordal complexes are weakly-ridge-chordal. (This expands Bigdeli, Yazdan-Pour and Zaare-Nahandi's work on ridge-chordality.)

math.CO↗

Zero-Flow Two-Sample Tests

Motivated by the success of modern flow-based generative models in modeling complex data, we study two-sample testing through the lens of flow-based methods. We propose the Zero-Flow Two-Sample Test (ZF2ST), built on the zero-flow criterion, which characterizes distributional equality through a time-reversal antisymmetry of a learnable velocity field. We extend this criterion and further develop the Zero-Flow Discrepancy, an identifying discrepancy that controls the Wasserstein distance, and derive a variational representation in terms of a witness function. This representation naturally leads to a witness-based test whose power is governed by the signal-to-noise ratio (SNR), allowing direct power maximization for witness learning. ZF2ST learns the witness on one data split and performs testing on held-out samples, thereby maintaining Type-I error control and admitting a simple asymptotic null distribution. Experimentally, ZF2ST performs competitively across various synthetic and real-world benchmarks, while showing particularly strong performance in distinguishing image distributions from different sources.

cs.LG↗

EnsembleEGNN: Set-Based Graph Learning for Thermodynamic Ensembles of Cyclic Peptides

Molecular graph encoding often relies on a single, static structure, ignoring the thermodynamic ensemble of molecules that are present in solution. Here, we introduce EnsembleEGNN, a foundation model that encodes structural ensembles by processing individual conformers through shared equivariant graph neural network layers, pooled with a set attention block, to make property predictions from the whole ensemble. Pretrained on the CREMP cyclic peptide dataset using multi-task self-supervision, the model is trained to encode the conformational variability of each molecule. When predicting membrane permeability from the CycPeptMPDB benchmark, EnsembleEGNN achieves an $R^2$ of $0.477$ under random cross-validation, outperforming a sequence-only BERT baseline ($R^2=0.439$). This representation advantage persists under rigorous out-of-distribution Butina splits ($R^2=0.401$ versus $0.354$). Finally, a hybrid architecture co-training EnsembleEGNN with the BERT model achieves the highest overall accuracy across both random ($R^2=0.538$) and structural holdout evaluations ($R^2=0.444$). These results demonstrate that encoding conformational ensembles into latent representations improves predictions for properties governed by thermodynamics.

cs.LG↗

Effect of Al-Zn alloy wafer grain boundary diffusion on the magnetism and microstructure of sintered NdFeB magnets

This study investigates Al-Zn grain boundary diffusion (GBD) treatment on sintered Nd-Fe-B magnets using $Al_{80}Zn_{20}$ alloy sheets as the diffusion source. The alloy sheets were placed at both ends of cylindrical samples and diffusion-annealed at 900$^\circ$C and 700$^\circ$C for 7 hours under vacuum ($\leq5\times10^{-3}$ Pa), followed by tempering at 500$^\circ$C for 2 hours. Magnetic measurements show that coercivity increases from 951.5kA/m in the untreated sample to 1158.2kA/m at 900$^\circ$C (a gain of 206.7kA/m, 21.7\%) and to 1039.6kA/m at 700$^\circ$C (a gain of 88.1kA/m, 9.3\%), while remanence declines modestly from 1282mT to 1256mT after the high-temperature treatment. Scanning electron microscopy (SEM), energy-dispersive X-ray spectroscopy (EDS), and X-ray diffractometer (XRD) analyses reveal that the 900$^\circ$C treatment produces a thinner, more continuous grain boundary phase and a distinct core-shell structure around the main-phase grains. EDS mapping shows that Al preferentially enriches the shell region of the $Nd_2Fe_{14}B$ grains, while Zn predominantly resides in the grain boundary phase, where it lowers the melting point of the intergranular phase and improves its fluidity. XRD shows no detectable secondary phases, while the combined XRD, EDS results, and lattice expansion are consistent with possible limited Al incorporation into the main-phase lattice. Verified by computational analysis, the coercivity enhancement is attributed to three synergistic factors: improved grain boundary decoupling, the formation of a high-anisotropy shell layer that strengthens domain-wall pinning, and the smoothing of grain edges to suppress reverse-domain nucleation.

cond-mat.mtrl-sci↗

Group Preference Collapse in Personalized Multimodal Large Language Models

Personalized multimodal large language models (MLLMs) aim to generate user-specific responses, but existing methods mainly rely on profile-level information and overlook diverse user preferences. We identify group preference collapse, where multi-user personalized MLLMs become insensitive to individual preferences and drift toward dominant population-level choices due to suppressed preference signals and unreliable preference use during generation. We propose PrefMoE, a preference-centric framework that separates stable profile information from preference-related representations. PrefMoE decomposes preferences into shared prototypes and personalized residuals, preserves individualized residuals with imbalance-aware learning, counterfactual pseudo-user augmentation, and residual decorrelation, and routes profile and preference factors through separate LoRA adaptation paths. Experiments across multiple MLLM backbones show that PrefMoE improves preference-sensitive personalization while substantially reducing preference collapse. Project page: https://prefmoe.github.io/.

cs.AI↗