Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 235 records · Page 13Linked to original sources

VectorMaton: Efficient Vector Search with Pattern Constraints via an Enhanced Suffix Automaton

Approximate nearest neighbor search (ANNS) has become a cornerstone in modern vector database systems. Given a query vector, ANNS retrieves the closest vectors from a set of base vectors. In real-world applications, vectors are often accompanied by additional information, such as sequences or structured attributes, motivating the need for fine-grained vector search with constraints on this auxiliary data. Existing methods support attribute-based filtering or range-based filtering on categorical and numerical attributes, but they do not support pattern predicates over sequence attributes. In relational databases, predicates such as LIKE and CONTAINS are fundamental operators for filtering records based on substring patterns. As vector databases increasingly adopt SQL-style query interfaces, enabling pattern predicates over sequence attributes (e.g., texts and biological sequences) alongside vector similarity search becomes essential. In this paper, we formulate a novel problem: given a set of vectors each associated with a sequence, retrieve the nearest vectors whose sequences contain a given query pattern. To address this challenge, we propose VectorMaton, an automaton-based index that integrates pattern filtering with efficient vector search, while maintaining an index size comparable to the dataset size. Extensive experiments on real-world datasets demonstrate that VectorMaton consistently outperforms all baselines, achieving up to 10x higher query throughput at the same accuracy and up to 18x reduction in index size.

cs.DB↗

Hadronic description of nuclear matter and neutron star properties

The composition of the neutron star is one of the most fundamental and long-standing problems in nuclear- and astro-physics. The known properties of nuclear matter, together with the astronomical observations, impose the stringent and interconnected constraints on the theoretical descriptions. In this work, by using the most general quantum hadrodynamics model including $σ, ω, ρ$ and $a_0$ in addition to nucleons, and performing a Bayesian joint analysis of experimental nuclear matter data, the heavy-ion flow pressure, and astrophysical observations including the mass-radius inferences of PSR J0740+6620, PSR J0437$-$4715, PSR J0614$-$3329 and HESS J1731$-$347 and the tidal posterior of GW170817, we point out that the nuclear matter made of only hadrons can provide a unified description of nuclear matter properties and astrophysical observations.In addition, we find that the speed of sound in the GQHD develops a non-monotonic, peak-like structure which is absent in the Walecka-type models TM1, NL3 and FSU-$\delta6.7$. This soft-to-stiff transition, accompanied by a pronounced softening of the symmetry energy, results in small size intermediate mass neutron stars, $R_{1.4}\simeq (11.1-11.4)$~km, together with the maximum mass $(2.2-2.3)M_\odot$, and, to our knowledge, has not been found before in the Walecka-type relativistic mean-field models. What we find here indicate that the sequential measurement of neutron star mass and radius by the next generation facilities, especially that of the intermediate mass neutron stars, is crucial for distinguishing the pure nucleonic stars from the hybrid ones.

nucl-th↗

MedFeat: Model-Aware and Explainability-Driven Feature Engineering with LLMs for Tabular Prediction

In clinical tabular prediction, classical machine learning models with feature engineering often outperform neural methods. LLMs are increasingly used to automate this process, acting as domain experts that propose diverse feature transformations to boost downstream performance. However, the feature generation process of existing LLM-based methods is agnostic to the downstream learner: the LLM receives no signal about which features currently drive predictions or where the model's representational capacity falls short, so proposals are neither targeted to promising regions of the feature space nor tailored to the learner's inductive bias. This shortcoming is amplified in healthcare data, which simultaneously exhibits class imbalance, heterogeneous feature spaces, and strict interpretability requirements. In this paper, we propose MedFeat, the first feature engineering framework inspired by the workflow of machine learning practitioners, leveraging model-awareness and feature importance signals to iteratively guide feature discovery for clinical tabular learning. We evaluate MedFeat on a broad range of challenging real-world clinical tasks and show that it statistically significantly outperforms state-of-the-art baselines, with an average F1 improvement of more than 10% over the baseline across models with distinct inductive biases.

cs.LG↗

The Alignment Flywheel: A Governance-Centric Hybrid MAS for Architecture-Agnostic Safety

Multi-agent systems provide mature abstractions for role decomposition, coordination, and normative governance, but increasingly capable learned components make post-deployment safety harder to inspect, audit, and update. When safety behavior is absorbed into a decision component, narrow failures may require retraining or rollback of the full component. This instantiates our vision of the Alignment Flywheel as a governance-centric hybrid MAS architecture that decouples decision generation from safety governance. We denote the agent or policy that generates candidate trajectories as the Proposer; it passes its output to a governed Safety Oracle stack, which returns safety scores, prediction uncertainty, audit coverage uncertainty, and evidence hooks through a stable interface. An Enforcement layer applies explicit risk policy at runtime. Around this loop, a governance MAS performs monitoring, red-teaming, verification, triage, refinement, and versioned release management. The central engineering principle is patch locality: many newly observed safety failures can be mitigated through small governance batches for the Oracle stack and its audit state rather than by retraining or retracting the Proposer. The architecture is implementation-agnostic with respect to both Proposer and Oracle. It defines the roles, artifacts, protocols, and release semantics needed for runtime gating, audit intake, signed updates, staged rollout, and rollback. We demonstrate executability in two scenarios: a learned spatial Oracle patched through regression-checked governance updates, and a clinical GenAI proxy setting illustrating structured norms, escalation, and audit coverage. Our implementation code and documentation are available open source at https://github.com/decide-ugent/Alignment-Flywheel.

cs.MA↗

Dinaturality for Double Categories

In this paper we extend the concept of dinaturality to the setting of double categories. We introduce the dinatural versions of double categorical transformations and modifications, and show that ordinary natural transformations and modifications correspond to dinatural ones between dummy functors. We prove the double categorical generalizations of some classical theorems about dinatural transformations for 1-categories, and extend the surface diagram calculus for the locally cubical Gray category of small double categories to include dinatural constructions. In our motivating example of dinatural constructions we reconstruct the double categorical formulation of mates in terms of dinatural transformations and characterize extranaturality for 1-categories using dinaturality for double categories.

math.CT↗

Momentum-projected hadron entanglement from lattice-QCD replica correlators

We define a finite-volume lattice-QCD observable for the vacuum-subtracted spatial Rényi response of a source--sink-prepared, momentum-projected hadron. At fixed regulator, integer replica index $n>1$, spatial region and gauge-theory cut prescription, double-sided state projection expresses this response as a ratio of replicated and ordinary hadron correlators. The first numerical target is the $n=2$ ratio; its proposed $L^{-3}$ scaling at fixed physical region is to be tested with matched volumes. We implement the measurement in quenched SU(3) for a charged pseudoscalar meson with heavy degenerate Wilson valence quarks, $am_{\rm eff}\simeq2.07$, on two $16^3\times32$ sheets with a central-plaquette cut. Five finite balls and empty/full controls are measured at nine source/sink placements. Their finite-ball means increase with radius at every placement. At the reference placement the three largest responses are 1.6--2.5 times an ideal single-quasiparticle occupancy reference. A full-covariance fit $Δ_{\rm cp}=κ_{\rm occ}S_2^{\rm occ}$ gives $κ_{\rm occ}=1.68\pm0.12$; radius-window and full-control comparisons give values spanning $1.49$--$1.73$. An exact pure-state quark--antiquark pair model interpolates, in the dilute limit, between the unit-occupancy reference ($κ_{\rm occ}=1$) for an unresolved meson and the double-occupancy reference ($κ_{\rm occ}=2$) for a separately resolved quark and antiquark, and yields a boundary-area correction governed by their separation.

hep-ph↗

FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering

Understanding long videos is crucial for embodied intelligent agents, as their performance depends on effectively accumulating and using long-horizon perceptual memories. Multimodal large language models (MLLMs) are increasingly used for long-video understanding, but their performance degrades and inference time increases as more frames are provided. Therefore, selecting informative keyframes is essential for efficient question answering over long videos. In this work, we develop FocusGraph, a framework for keyframe selection in egocentric long-video question answering. It includes a lightweight Scene-Graph LLM Selector that identifies query-relevant clips from compact graph-based captions, avoiding the need to process raw frame sequences at question time. From these clips, we extract keyframes using Patch-wise Sparse-Flow Retention (PSFR), an offline program-evolved method with no learned parameters at inference time, before passing them to an MLLM for answer generation. FocusGraph achieves state-of-the-art performance on FindingDory and HourVideo while reducing question-time inference cost compared with existing approaches.

cs.CV↗

A Multi-Fidelity Tensor Emulator for Spatiotemporal Outputs: Emulation of Arctic Sea-Ice Dynamics

Numerical models are widely used to simulate the Earth system, but they are computationally expensive and often depend on many uncertain input parameters. Calibration and uncertainty quantification require running them across many input configurations, so their effective use is costly. Statistical emulation provides a practical alternative for efficiently exploring model behavior. We are motivated by the Arctic sea-ice component of the Energy Exascale Earth System Model (MPAS-Seaice), which generates large spatiotemporal outputs at multiple spatial resolutions, with high-resolution (or high-fidelity, HF) simulations being more accurate but computationally more expensive than lower-resolution (low-fidelity, LF) simulations. Multi-fidelity (MF) emulation integrates information across resolutions to construct efficient and accurate surrogate models, yet existing approaches struggle to scale to large spatiotemporal data. We develop a MF emulator that combines tensor decomposition for dimensionality reduction, Gaussian process priors for flexible function approximation, and an additive discrepancy model to capture systematic differences between LF and HF data. The proposed framework scales to complex spatiotemporal fields while keeping predictions accurate and uncertainty well calibrated. In the MPAS-Seaice analysis, it attains the lowest prediction error of all emulators considered, including separable-covariance and neural-network approaches, especially throughout the summer melt season, when sea-ice concentration varies most across model settings and matters most scientifically, and best reproduces the spatiotemporal dependence of the output. By combining many inexpensive low-resolution runs with a few expensive high-resolution ones, the emulator substantially reduces the simulation cost of exploring model behavior, and suits large-scale studies involving complex physical models.

stat.ME↗

Strongly interacting singlet scalar dark matter during reheating

We revisit the singlet scalar dark matter model in the presence of a non-standard cosmological history prior to radiation domination. We focus on the regime in which the relic abundance is set by 4-to-2 self-annihilations while the dark and visible sectors remain in kinetic equilibrium, i.e. the standard strongly interacting massive particle (SIMP) framework. In the conventional radiation-dominated cosmology, this realization is not viable, as it requires sub-MeV masses and large quartic couplings in tension with bounds on dark matter self-interactions. We show that this conclusion is significantly modified if freeze-out occurs during non-standard cosmological eras. The altered Hubble expansion rate and the possible non-conservation of the standard model entropy change the freeze-out dynamics, allowing the observed relic density to be achieved with perturbative couplings and consistent with astrophysical constraints. We determine the region where SIMP production dominates over the WIMP mechanism and confront the viable parameter space with current and future direct detection and collider bounds.

hep-ph↗

Optimality as a Generative Principle for Network Structure

Most real-world networks have evolved or been engineered to optimize some function, yet a unified framework for studying optimal networks across domains is lacking. We introduce GradNet, an AI-enabled framework that treats network topology as a continuously differentiable object, enabling the design and study of networks that optimize structural and dynamical objectives amenable to automatic differentiation under realistic constraints. We derive general optimality conditions, including an equimarginal principle for linear budgets, that make optimized networks analytically tractable. Canonical network features emerge spontaneously from constrained optimization: maximizing Kuramoto synchronization under coupling budgets yields sparse, bipartite, frequency-disassortative networks; minimizing social tension in opinion dynamics reproduces the factional split in Zachary's karate club; and maximizing communication capacity in spatial quantum networks under distance-dependent costs recovers minimum spanning trees. GradNet thus serves both as a network design tool scalable beyond $10^5$ nodes and as a scientific probe of structure-function relationships.

physics.soc-ph↗

Vibrational strong coupling influences product selectivity in a model for post transition state bifurcation reactions

In this study we explore the possibility of modulating product selectivity (branching ratios) in post transition state bifurcation (PTSB) reactions via vibrational strong coupling (VSC) to an optical cavity. Detailed classical and quantum dynamical calculations on a model potential reveal that the branching ratio can be enhanced by nearly a factor of two under VSC conditions. Interestingly, upon altering the shape of the potential, we find a switch in the cavity frequency at which maximum enhancement in selectivity is observed. Apart from emphasizing the role of both cavity-system and intramolecular energy transfer to the observed enhancements, we highlight the complexity of the VSC mechanism in terms of the choice of the cavity frequency vis--à--vis the various molecular mode frequencies. Our work shows that, in principle, cavity quantum electrodynamics can reshape dynamical outcomes in reactions with complex potential energy landscapes.

physics.chem-ph↗

Multi-Task Anti-Causal Learning for Reconstructing Urban Events from Residents' Reports

Many real-world machine learning tasks are anti-causal: they require inferring latent causes from observed effects. In practice, we often face multiple related tasks where the structural dependencies are a hybrid of task-invariant and task-specific mechanisms. We propose Multi-Task Anti-Causal learning (MTAC), a framework for estimating causes from outcomes and confounders by explicitly exploiting such cross-task invariances. MTAC learns a structural equation model (SEM) that factorizes the outcome-generation process into (i) a task-invariant mechanism and (ii) task-specific mechanisms via a shared backbone with task-specific deviations. Building on the learned forward model, MTAC performs maximum A posteriori (MAP) based inference to reconstruct causes by jointly optimizing latent mechanism variables and cause magnitudes under the learned structural model. We evaluate MTAC on the application of urban event reconstruction from resident reports, spanning three tasks: parking violations, abandoned properties, and unsanitary conditions. On real-world data collected from Manhattan and Newark, MTAC consistently improves reconstruction accuracy over strong baselines, achieving up to 33.04\% MAE reduction and demonstrating the benefits of learning transferable mechanisms across tasks.

cs.LG↗

Cross-Context Review: Improving LLM Output Quality by Separating Production and Review Sessions

Large language models struggle to catch errors in their own outputs when the review happens in the same session that produced them. This paper introduces Cross-Context Review (CCR), a straightforward method where the review is conducted in a fresh session with no access to the production conversation history. We ran a controlled experiment: 30 artifacts (code, technical documents, presentation scripts) with 150 injected errors, tested under four review conditions -- same-session Self-Review (SR), repeated Self-Review (SR2), context-aware Subagent Review (SA), and Cross-Context Review (CCR). The central result is that a second review helps only when it happens in a fresh session: CCR (F1 28.6%) outperforms a second review in the same session (SR2, 21.7%) robustly, both in the first run (paired t, p<0.001) and in the three-run average (Holm-adjusted p=0.004). This version updates the broader comparisons. Averaged across runs, and excluding one SR run whose records could not be verified, CCR is not significantly ahead of context-aware subagent review (SA, 23.8%; p=0.057) or of a single same-session review (SR, 27.1%; p=0.26); the first version's advantages over these two baselines came from run 1. CCR needs no infrastructure and costs one extra session.

cs.CL↗

The COTe score: A decomposable framework for evaluating Document Layout Analysis models

Document Layout Analysis (DLA) is the process by which a page is parsed into meaningful elements, often using machine learning models. Typically, the quality of a model is judged using general machine vision metrics such as IoU, F1 or mAP. However, these metrics are designed for images that are 2D projections of 3D space, not for the natively 2D imagery of printed media. This discrepancy can result in misleading or uninformative interpretation of model performance. To encourage more robust, comparable, and nuanced DLA, we introduce: The Structural Semantic Unit (SSU), a relational labelling approach that shifts the focus from the physical to the semantic structure of the content; and the Coverage, Overlap, Trespass, and Excess (COTe) score, a decomposable metric for measuring page parsing quality. We demonstrate the value of these methods through case studies and by evaluating 5 common DLA models on 3 DLA datasets. We show that the COTe score is more informative than traditional metrics and reveals distinct failure modes across models, such as breaching semantic boundaries or repeatedly parsing the same region. We find that, under granularity differences between model and ground truth, the COTe score is substantially more robust than the F1. Even in the worst case, comparing character-level predictions against paragraph-level ground truth with otherwise perfect parsing, COTe returns 0.68 where F1 returns 0. Notably, we find that, on real datasets, the COTe's granularity robustness largely holds even without explicit SSU labelling, reducing the barrier to entry. Finally, we release an SSU labelled dataset and a Python library for applying COTe in DLA projects.

cs.CV↗

Twin-peaked gravitational wave signals from a $\mathcal{Z}_2$ phase transition

We compute the gravitational wave spectrum from a phase transition associated with the spontaneous breaking of a $\mathcal{Z}_2^{\rm DW}$ symmetry. If the transition is second-order, the only source of gravitational waves is the annihilation of domain walls (biased by quantum gravity) formed after this breaking. However, if the transition is first-order, this yields a twin-peaked signal from both the transition itself and the biased domain wall annihilation. Both scenarios originate when a scalar singlet odd under the $\mathcal{Z}_2^{\rm DW}$ obtains a non-zero vacuum expectation value. An additional $\mathcal{Z}_2^{\rm DM}$ odd scalar doublet strengthens the transition by keeping the singlet scalar in thermal equilibrium with the Standard Model plasma at early times. Additionally, the same scalar doublet produces fermionic dark matter via freeze-in, matching observed dark matter relic density.

hep-ph↗

Overlapping Covariance Intersection: Fusion with Partial Structural Knowledge of Correlation from Multiple Sources

Emerging large-scale engineering systems rely on distributed fusion for situational awareness, where agents combine noisy local sensor measurements with exchanged information to obtain fused estimates. However, at the sheer scale of these systems, tracking cross-correlations becomes infeasible, preventing the use of optimal filters. Covariance intersection (CI) methods address fusion problems with unknown correlations by minimizing worst-case uncertainty based on available information. Existing CI extensions exploit limited correlation knowledge but cannot incorporate structural knowledge of correlation from multiple sources, which naturally arises in distributed fusion problems. This paper introduces Overlapping Covariance Intersection (OCI), a generalized CI framework that accommodates this novel information structure. We formalize the OCI problem and establish necessary and sufficient conditions for feasibility. We show that a family-optimal solution can be computed efficiently via semidefinite programming, enabling real-time implementation. The proposed tools enable improved fusion performance for large-scale systems while retaining robustness to unknown correlations.

eess.SY↗

Do Vision Language Models Understand Human Engagement in Games?

Inferring human engagement from gameplay video is important for game design and player-experience research, yet it remains unclear whether vision--language models (VLMs) can infer such latent psychological states from visual cues alone. Using the GameVibe Few-Shot dataset across nine first-person shooter games, we evaluate three VLMs under six prompting strategies, including zero-shot prediction, theory-guided prompts grounded in Flow, GameFlow, Self-Determination Theory, and MDA, and retrieval-augmented prompting. We consider both pointwise engagement prediction and pairwise prediction of engagement change between consecutive windows. Results show that zero-shot VLM predictions are generally weak and often fail to outperform simple per-game majority-class baselines. Memory- or retrieval-augmented prompting improves pointwise prediction in some settings, whereas pairwise prediction remains consistently difficult across strategies. Theory-guided prompting alone does not reliably help and can instead reinforce surface-level shortcuts. These findings suggest a perception--understanding gap in current VLMs: although they can recognize visible gameplay cues, they still struggle to robustly infer human engagement across games.

cs.CV↗

A bilinear inverse problem with forward operator inaccuracy applied to neonatal atlas-based diffuse optical tomography

In this work, we assume to have a set of candidate forward operator matrices and suggest principal component analysis for modeling their variation from the mean. We use this principal component representation to adapt a bilinear inverse-problem formulation to forward-operator inaccuracy in neonatal atlas-based diffuse optical tomography and present two optimization algorithms, as well as Gibbs sampling and a version of the Bayesian approximation error method, for approximately solving the resulting problem. We apply the algorithms to account for the inaccuracy that is present in the sensitivity profiles or Jacobian matrices in diffuse optical tomography when an atlas-based model of the head anatomy is used instead of the subject's own anatomical model in neonates over a wide range of gestational ages (29--44 weeks). We report visual and numerical improvements in the spatial localization and contrast-to-noise-ratio in reconstructions of simulated hemodynamic activity.

math.NA↗