Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Recovering Governing Dynamics from Distributed Observations via Exact Spline Merging

Scientific observations are frequently distributed across locations, time periods, and institutions. Combining such observations into a continuous, differentiable field enables recovering governing physical parameters from its derivatives. This paper makes two contributions in this setting. First, the established additive structure of fixed-basis ridge-regression statistics is applied to tensor-product spline fields: each data holder computes a local Gram matrix and moment vector, and the merged solution is mathematically identical to centralized fitting, with no raw data shared and no iterative synchronization. This property is specific to the fixed-feature squared-error setting; the present derivation does not establish an analogous guarantee for general jointly trained multilayer networks. Second, a complete pipeline connects distributed observations to physical parameter inference through field reconstruction, derivative extraction, and linear regression. The pipeline is validated on four PDEs: diffusion, wave, heat-with-source, and the nonlinear viscous Burgers equation, recovering governing parameters to sub-percent accuracy in the linear cases and 5\% for Burgers. In all cases, distributed merging introduces zero degradation relative to centralized fitting. Synthetic experiments validate parameter recovery; application to 41 years of NOAA sea-surface temperature data validates field reconstruction and aggregation equivalence on real spatiotemporal observations. Source code to reproduce all experiments is available at https://github.com/NAVEENMN/splinemerge.

cs.LG↗

A Vision-Language Foundation Model for Precise and Comprehensive Brain Tumor Diagnosis from Preoperative Multimodal Data

We developed BrainVLM to classify all 12 World Health Organization (WHO) 2021 brain tumor types. BrainVLM integrates an uncertainty quantification strategy to indicate prediction reliability and a module for generating radiology reports to elucidate the clinical rationale. BrainVLM was trained on multi-modal data (MRI scans, demographics, and radiology reports) from 40,043 individuals. It was validated on 5,211 patients with pathologically confirmed brain tumors, including 3,877 held-out patients from the primary hospital and 1,334 patients from 11 independent hospitals. We further conducted two proof-of-concept studies to validate its clinical utility in AI-clinician workflows: 1) a blinded multireader study where 12 neuroradiologists across varying experience levels interpreted 248 retrospective cases with or without AI assistance, and 2) a real-world prospective study in which 1,009 patients were independently and blindly assessed by BrainVLM and radiologists before surgery. Additionally, we demonstrated BrainVLM's utility in preoperative molecular subgroup prediction for adult-type diffuse gliomas, using a multi-center cohort of 632 patients. In primary evaluation, BrainVLM achieved an area under the curve (macro-AUC) of 0.85 (95% CI: 0.84-0.86), and an F1 score of 0.82 (95% CI: 0.81-0.83), surpassing neuroradiologists (F1 = 0.80 (95% CI: 0.79-0.81)). In external validation across 11 centers, BrainVLM achieved an AUC = 0.80 (95% CI: 0.79-0.82) and F1 = 0.75 (95% CI: 0.73-0.78), compared with F1 = 0.71 (95% CI: 0.69-0.73) for neuroradiologists. In prospective real-world evaluation, BrainVLM maintained performance comparable to neuroradiologists. The BrainVLM project page is available at https://hku-healthai.github.io/brainvlm_project.github.io/.

cs.CV↗

A Weighted Kernel Method for Approximation that Adapts to Learned Multivariable Structure

Approximating the input-output behavior of a multivariable black-box function from limited data is challenging when blind to the importance of its inputs and their interactions. We introduce total sensitivity kernels (TSKs), a method based on families of weighted ANOVA kernels that learn and adapt to this multivariable structure. TSKs parameterize the weights on each multivariable component of the target function by factors for each input. We propose learning these factors directly from function evaluations by selecting the reproducing kernel Hilbert space (RKHS) in which the target function has minimum norm. Under suitable conditions, we show that this norm-minimization problem admits a unique solution, and we establish consistency of a finite-data formulation based on minimum-norm interpolation. The learned TSK factors characterize the participation of individual inputs across interactions and main effects, providing a kernel-dependent notion of input sensitivity related to total Sobol indices. Numerical experiments demonstrate that adapting the kernel to learned multivariable structure can substantially improve approximation accuracy over a standard product kernel.

cs.LG↗

Information Geometric Self-Organization at the Edge of Stability in High-Capacity Kernel Associative Memories

High-capacity associative memories based on Kernel Logistic Regression (KLR) exhibit exceptional storage capabilities and robustness. Previous empirical studies identified a hyperparameter regime, the "Ridge of Optimization," where attractor stability is maximized. However, the geometric nature of this regime and the optimization dynamics required to reach it have remained unclear. In this paper, we investigate the static geometry of the parameter space and the learning trajectory of Gradient Descent (GD) in KLR-trained Hopfield networks. Using the eigenvalue spectrum of the Hessian, we reveal that the Ridge corresponds to a phase boundary located adjacent to a rank-1 spectral collapse, acting as a geometric singularity where the principal curvature is massively amplified. Furthermore, we demonstrate that the learning dynamics exhibit a transient self-stabilizing behavior driven by the Edge of Stability (EoS) phenomenon. Rather than seeking flat regions, the network parameters are driven toward a state where the local curvature dynamically equilibrates near the stability limit dictated by the learning rate, allowing the optimization to survive the initial instability. We provide analytical derivations for both the rank-1 asymptotic collapse and the dynamic feedback loop governing this equilibration. These findings suggest that optimal, high-capacity memory representations are not formed in flat minima, but are dynamically sculpted at the highly curved boundaries of geometric singularities.

cs.LG↗

Beyond Quadratic Loss: The Stability Phase Diagram of Adam

Loss spikes are recurrent instabilities in neural-network training and can arise from multiple mechanisms. For Adam in particular, macroscopic loss spikes have been linked to optimizer dynamics, yet how its two momentum timescales govern them remains unclear. We investigate this dependence by mapping training dynamics across the $(β_1,β_2)$ plane. Across a range of model--task settings, an approximately linear boundary, $1-β_2=C(1-β_1)$, separates spiky from non-spiky dynamics, whereas a one-dimensional quadratic loss produces approximately cubic slope. A one-dimensional superquadratic loss $L(x)\propto|x|^n$ recovers the near-linear scaling and links the boundary coefficient to the effective loss exponent $n$. We further show that confident cross-entropy losses develop a core--wall landscape comprising a narrow quadratic core followed by a steep wall, which produces effective superquadratic behavior at the scale of an optimizer update. Together, these results connect Adam loss spikes to both the mismatch between momentum timescales and finite-scale superquadratic loss geometry beyond the Hessian.

cs.LG↗

The Reciprocal Problem on Weighted Bergman Spaces

We study the reciprocal problem for weighted Bergman spaces: if $f\in A_α^p(\Bn)$ and $\inf_{\Bn}|f|>0$, does it follow that $1/f\in A_α^p(\Bn)$? We determine the range of parameters for which the answer is affirmative. In particular, we prove that functions in $A_α^p(\D)$ have the reciprocal property for all $α\in \R$ and $p\geq 1,$ where $\D$ denotes the unit disk in $\C.$ Moreover, functions in $A_α^p(\Bn)$ have the reciprocal property for all $α\in \R$ when $n=2$ and $p=2.$ In addition, we resolve the reciprocal problem in the three-dimensional Drury--Arveson space $H_3^2$ and obtain an equivalent condition in the four-dimensional space $H_4^2$. We also settle the reciprocal problem for general $A_α^p(\Bn)$ under some additional conditions.

math.CV↗

Revisiting the Brunner-Munzel test from the viewpoint of local linear approximation

The Brunner-Munzel (BM) test is a nonparametric test for two independent samples that evaluates whether observations from one group tend to be greater than observations from another group, or vice versa. The BM test has a broader scope of application than the Mann-Whitney $U$ test because it does not assume equal variances between the two groups. However, the meaning of the BM test statistic is difficult to understand intuitively, which may be one of the factors hindering the widespread use of the BM test. To alleviate this problem, in this paper, I introduce an alternative interpretation of the BM test statistic from the viewpoint of local linear approximation. It is shown that the variance estimator for the sample stochastic superiority used in the BM test can be derived using local linear approximation, in which the influence of each observation on the sample stochastic superiority is assumed to be additive. This simple interpretation will help practitioners decide to use the BM test without hesitation.

math.ST↗

Revisiting the Transfinite Christensen-Pedersen Argument

Christensen and Pedersen proved that every properly infinite $\mathrm{AW}^*$-algebra is monotone sequentially complete, and Saitô and Wright developed a transfinite form of their dilation argument. We revisit the transfinite construction using normality of $\mathrm{AW}^*$-algebras. Normality simplifies the limit stages by turning suprema into compressions of joins, so the construction only needs a supply of fresh orthogonal projections large enough to contain the supports of the summands at successor stages. We use this simplified proof to show that a $*$-homomorphism between $\mathrm{AW}^*$-algebras that preserves only the joins needed to encode such a sum preserves the sum itself. We also use it to deduce order-continuity facts about $κ$-join-preserving $*$-homomorphisms. We also show that a finite $\mathrm{AW}^*$-algebra has suprema for all bounded positive families whose supports have bounded total center-valued dimension.

math.OA↗

Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

Massive KV caches can cause severe memory-bandwidth bottlenecks during long-context decoding. Sparse attention methods mitigate this problem, but often drop necessary context, leading to quality degradation. We introduce \textbf{Elastic Threshold Attention (ETA)}, an end-to-end trainable architecture that rivals dense model quality under hardware-aligned block-sparse decoding. ETA predicts dynamic, contextual thresholds directly from query representations, adjusting context retention depending on the task at hand. To learn this policy from scratch while enabling near lossless KV cache pruning at inference time, we filter attention logits through a shifted SiLU gate during training. We show theoretically and empirically that this creates a smooth, near-uniform attention floor that neutralizes sub-threshold value contributions while simultaneously causing localized attention sinks on initial tokens to disappear. To materialize these advantages, we implement a fused inference-time kernel in Triton that screens KV blocks in $O(1)$ time using cached geometric-probabilistic bounds. Across language modeling, reasoning, and RULER benchmarks, our 1.45B ETA model matches or exceeds dense quality, outperforming alternative fast decoding methods (Quest, H$_2$O, NSA) while achieving higher sparsity levels. Our kernel also achieves up to $2.15\times$ end-to-end speedups over FlashAttention-2 at context lengths of up to $512$K tokens. Finally, we introduce an offline calibration algorithm for domain-specific deployments that freezes per-head constant thresholds, cutting attention compute by an additional 27\% at no quality cost.

cs.LG↗

Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining

Conventional tokenizers represent text as characters or statistically derived subwords, overlooking the internal phonological structure of syllables and often requiring large vocabularies. We introduce \textbf{Phonemic Tokenizer}, a linguistically motivated tokenizer for Vietnamese and Chinese that converts each syllable into IPA and factorizes it into three phonological components: onset, rime, and tone. The three components jointly occupy one contextual position, preserving syllable-level sequence length while enabling representation sharing across phonologically related syllables. Non-phonological and unsupported units are handled through character-level fallback. This deterministic design requires no corpus-dependent vocabulary learning and yields vocabularies of only 112 entries for Chinese and 256 for Vietnamese. Intrinsic evaluation shows that the tokenizer achieves substantially higher Rényi efficiency in both languages, represents every entry in a standard Vietnamese syllable dictionary with a Fertility of exactly one, and generally produces shorter Vietnamese sequences than existing pretrained tokenizers. We further instantiate the tokenizer in \textbf{PhonemicBERT}, which combines factorized component embeddings and reconstructs complete masked syllables using three prediction heads. Under a controlled Chinese pretraining setup, PhonemicBERT-Zh is competitive with or outperforms character, subword, and SubChar alternatives across diverse language-understanding tasks. PhonemicBERT-Vi also achieves competitive or superior results to established Vietnamese and multilingual pretrained models. These results establish phonemic factorization as a compact, efficient, and interpretable alternative to atomic and statistically segmented text representations.

cs.CL↗

An involution on Dyck paths in the region $\defc\le\min(\area,\dinv)$ that interchanges area and dinv

We construct an explicit involution on Dyck paths satisfying $\defc\le\min(\area,\dinv)$ that interchanges area and dinv. The construction relies on a number of new combinatorial objects developed here and two main external tools: the dual Dyck insertion of \cite{Hawkes26} and the Garsia--Milne involution principle~\cite{GarsiaMilne81}, as formulated by Doyle~\cite{Doyle19}. In addition, we give an involution on Dyck paths with $\defc \le n-3$ that interchanges area and dinv. This last task relies on work of Loehr and Warrington~\cite{LoehrWarrington09}.

math.CO↗

Generalized staircase partitions and Macdonald principal specializations

Generalized staircase partitions are obtained by replacing each box of an ordinary staircase with a fixed rectangle. We determine the unique shortest horizontal-strip sequence between consecutive generalized staircases and its conjugate vertical-strip sequence. For monic Macdonald polynomials, we derive explicit finite principal-specialization ratios along both sequences. For arbitrary nested partitions, coefficientwise nonnegativity after removal of the monomial factor forces independence of $q$. Along the horizontal sequence, this nonnegativity is characterized by rectangular complementation, apart from the one-column case. Along the vertical sequence, it occurs at the smallest admissible number of variables, apart from the initial column case. Endpoint ratios yield a triangular product formula and a centered product with exchange, reciprocity, and inversion identities. Hall-Littlewood, Jack, and Schur specializations give Gaussian-polynomial, finite-product, and tableau formulas, respectively.

math.CO↗

Rethinking Human-Aligned Evaluation: An Analysis of Semantic Metrics Beyond WER

Word Error Rate (WER), the most commonly used metric for Automatic Speech Recognition (ASR), treats every lexical deviation from the reference as equally costly, regardless of whether it changes meaning. This raises the question: does WER actually track how humans judge ASR transcript quality? We introduce HATS-en, an English dataset for human-centered ASR evaluation. Using this dataset, we benchmark lexical metrics against several configurations of BERTScore and SemDist, varying the language model, layer, and pooling strategy. We find that WER agrees least with human judgment among all metrics tested, that the best-performing SemDist configurations achieve the highest overall agreement, ahead of CER and BERTScore, and that no single model is best across settings. CER, despite its simplicity and low cost, remains remarkably close to these best configurations. In line with prior recommendations, our results support shifting ASR evaluation toward CER both for English and for morphosyllabic writing systems as it is a more interpretable and low-cost metric for what evaluation should actually capture, and using SemDist as a complementary evaluation.

cs.CL↗

High Resolution Grid-based Simulations of the Warm-Hot Intergalactic Medium

We present high-resolution cosmological hydrodynamic simulations of the Warm-Hot Intergalactic Medium (WHIM) using the GPU-optimized grid-based hydrodynamic code Kratos. Employing a uniform 4096^3 grid in a (100 h^-1 Mpc)^3 comoving volume, we achieve a spatial resolution of ~24.5 h^-1 kpc, sufficient to resolve the Jeans scale of gas at T ~ 10^4 K and n_H ~ 10^-3 to 10^-2 cm^-3. This calculation ranks among the largest grid-based cosmological hydrodynamic simulations performed to date. We find that ~23.4% of cosmic baryons reside in the WHIM phase (T = 10^5 - 10^7 K) at z = 0, significantly below the 40-50% found in earlier, lower-resolution simulations. Through a suite of lower-resolution simulations, we demonstrate that spatial resolution plays a pivotal role in determining the WHIM fraction: resolving gas near its Jeans scale allows it to reach higher densities, where enhanced radiative cooling transfers a substantial fraction of baryons out of the WHIM temperature range. The remaining WHIM resides predominantly in filaments and accretion-shock structures in the vicinity of halos, where hierarchical structure formation and ongoing hydrodynamic accretion provide continued shock heating. Synthetic observations of Lyman-alpha, O VI, O VII, and O VIII emission reveal distinct morphological and kinematic signatures, with O VI tracing filament-halo interfaces and the X-ray lines probing hotter gas associated with massive halos. These predictions underscore the importance of current and future missions such as XRISM, ATHENA, and HUBS for mapping WHIM thermodynamics and kinematics.

astro-ph.CO↗

Lyapunov stability of polynomial vector fields is undecidable

We show that there are integers $N$ and odd $D$ such that no algorithm can decide, from the rational coefficients of a homogeneous polynomial vector field $F$ in dimension $N$ of degree $D$, whether the origin is Lyapunov stable for $\dot Y=F(Y)$. This proves a conjecture of V. I. Arnold. We provide a Lean formalization of the proof. The smallest dimension $N$ for which we were able to prove undecidability is $N = 5$. An analogous undecidability result holds for global asymptotic stability and global exponential stability for non-homogenous vector fields. For dimension $N=2$ and homogenous vector fields, we establish decidability for all commonly used stability notions, building heavily on existing results.

math.DS↗

Boundary Kernel Rigidity and Infinite-Dimensional Complex Hyperbolic Representations

For n>=2, we prove a rigidity theorem for the Hermitian boundary kernels L_{t,s}(x,y)=|1- |^t exp(is arg(1- )), x != y, on S^{2n-1}, with L_{t,s}(x,x)=0, t>0, and s in [-1,1]. Every finite Gram matrix of L_{t,s} has positive index at most one if and only if 0<t<=1 and s=+-t. The proof combines Fourier analysis on boundary circles with restrictions to two orthogonal complex directions. The dimension threshold is sharp. By normalizing the Gram kernels of equivariant boundary maps, we show that the translation-length factor t and the signed Cartan factor s of every continuous irreducible representation from PU(n,1) to the holomorphic isometry group of infinite-dimensional complex hyperbolic space satisfy s=+-t. Combining this with the known endpoint rigidity, Monod's constructions, and Ruiz Stolowicz's complete-invariant theorem yields the classification: the holomorphic conjugacy classes are parametrized by (t,epsilon) in (0,1) x {+-1}. Allowing anti-holomorphic conjugacy identifies the two signs. Using Monod's automatic-continuity argument and the same boundary rigidity, we also classify all irreducible self-representations of the full isometry group of infinite-dimensional complex hyperbolic space. Up to conjugacy, these are precisely Monod's representations with parameter 0<t<=1. In particular, the canonical-orbit hypothesis in Monod's classification theorem is automatic.

math.GR↗

Probing Matter Unification through Higgs Physics at the LHC

We investigate the Higgs sector phenomenology of the minimal low scale quark-lepton unification theory. In addition to the Standard Model-like Higgs boson, the theory predicts a new CP-even neutral Higgs, a CP-odd neutral Higgs, and two charged Higgses, together with heavy right-handed neutrinos and leptoquarks. We analyze the scalar spectrum, the Yukawa structure responsible for charged-fermion masses, and the neutrino sector realized through an inverse-seesaw mechanism, emphasizing that the couplings of the new Higgs bosons are directly tied to the structure of matter unification and therefore lead to characteristic decay patterns dominated by third-generation fermions. We compute the branching ratios and production cross sections of the new Higgs states at the Large Hadron Collider (LHC) and assess their impact on present and future collider measurements. We find that current LHC data probe significant regions of parameter space for the neutral heavy Higgs bosons, with the strongest sensitivity arising from channels involving tau leptons, tops, and related final states, while the charged Higgs remains less constrained because of its electroweak production mechanism. Our results show that Higgs measurements provide a powerful and complementary probe of low-scale matter unification beyond direct leptoquark searches and open a new avenue for testing quark-lepton unification at present and future colliders, with discovery potential in high-luminosity LHC running.

hep-ph↗

Grounded Action Model: 3D Grounding as a Foundation for Robotics

Manipulation policies must know which objects matter and where they are, yet the pretrained backbones that current robot foundation models build on, from language in vision-language-action models (VLAs) to video generation in world-action models (WAMs), do not directly require this metric grounding, leaving it to be learned implicitly from robot demonstrations. We propose Grounded Action Models (GAMs), a new paradigm of robot foundation models built with 3D grounding. GAM can be conditioned using language, points, or box prompts, which are first transformed into a shared object-centric representation of the selected objects. This representation captures target-focused visual features and metric object geometry, which is mixed with robot state history through a multi-stream transformer to predict action chunks. Although GAMs can be run autonomously, they can also serve as a low-level controller that a high-level planner controls using its various input modalities, allowing for long-horizon and memory-dependent manipulation. On RoboTwin 2.0, GAM achieves an average success rate of 55.3% across 50 tasks (vs. 52.0% for Spatial Forcing), including 47.6% under scene randomization (vs. 30.4% for Abot-M0), with its action policy trained only on clean-scene demonstrations. On LIBERO-PRO, it achieves a state-of-the-art average success rate of 61% (vs. 53% for $π_{0.5}$) across 16 perturbation settings, with the largest gains when targets are relocated or newly designated. On two real robots, GAM retains 17/20 successes under visual shift on a bimanual YAM versus 4/20 for $π_{0.5}$, while its composition with a Molmo2 planner on a Franka achieves 64.7% ID and 49.8% OOD step completion on long-horizon and memory-dependent tasks.

cs.RO↗