Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 289 records · Page 16Linked to original sources

AEGIS: Anchor-Enforced Gradient Isolation for Knowledge-Preserving Vision-Language-Action Fine-Tuning

Fine-tuning pre-trained Vision-Language Models (VLMs) for robotic manipulation introduces a fundamental stability-plasticity dilemma: continuous flow-matching action experts backpropagate concentrated, low-rank regression gradients into transformer backbones trained on high-dimensional cross-entropy objectives. This cross-modal gradient asymmetry rapidly degrades pre-trained visual reasoning. Existing solutions either disconnect continuous gradient flow via stop-gradients or constrain updates via LoRA, which restricts update rank but remains directionally blind to semantic corruption; both typically rely on mixed-batch VQA co-training, doubling training compute. We introduce AEGIS (Anchor-Enforced Gradient Isolation System), a buffer-free, layer-wise orthogonal gradient projection framework enabling continuous flow-matching fine-tuning while isolating pre-trained representations from destructive parameter updates. Prior to training, AEGIS estimates per-layer Gaussian activation statistics from pre-training data as a static reference anchor. During fine-tuning, a closed-form Wasserstein-2 transport penalty generates an anchor-restoration gradient through the active computation graph. A sequential dual-backward pass applies layer-wise Gram-Schmidt orthogonalization, projecting task gradients onto the orthogonal complement of the restoration vector during directional conflict. We establish an exact energy preservation bound for layer-wise orthogonal projection, showing that AEGIS sheds only 0.62% of gradient energy empirically while halting cumulative feature drift. On PaliGemma2-3B fine-tuned on the LIBERO manipulation benchmark, AEGIS fully preserves pre-trained Visual Question Answering performance and baseline holdout loss while matching continuous action convergence, without replay buffers, teacher models, or co-training data.

cs.LG↗

Clinical Note Bloat Reduction for Efficient LLM Use

Background: Clinical notes contain extensive duplicated text from templates, copy-paste, and auto-populated fields ("note bloat"), diluting clinical signal, limiting longitudinal context, and increasing large language model (LLM) costs. Methods: TRACE removes note bloat using note-level EHR metadata to identify templated and copied content, with frequency-based de-duplication when metadata are unavailable. We evaluated TRACE using blinded physician span review and gold-standard templated-text annotations across four cohorts spanning liver transplant, obstetrics, and inpatient populations at multiple health systems (5.3M notes). We compared zero-shot LLMs and embedding-based classifiers using original and TRACE-processed notes for 20 information extraction tasks and prediction of 5-year survival, postpartum hemorrhage, and 30-day readmission. Results: Only 0.3-6.6% of removed text was flagged as author-generated; TRACE captured 86% of annotated templated characters. Information extraction F1 differences averaged by cohort ranged from -0.009 to +0.004; task-specific prediction F1 differences ranged from -0.011 to +0.018. Among 1,000 randomly sampled Stanford Health Care patients, TRACE reduced chart text by 47.3% (742.7M characters), averaging 220,167 fewer tokens per patient. Using 2024 encounter volumes at a large tertiary academic center and one query per encounter, projected three-year net savings ranged from $1.00M to $13.58M across evaluated model pricing schemes, including initial and annual TRACE processing costs. Conclusion: TRACE substantially reduces clinical note redundancy while preserving information extraction and prediction performance. Underused EHR metadata can reduce LLM inference costs, expand usable longitudinal context, and support scalable clinical AI.

cs.CY↗

Recursive determinantal framework for testing D-stability

The concept of matrix $D$-stability, introduced in 1958 by Arrow and McManus, is of major importance across a wide variety of applications in economic modeling, ecology, and control systems. However, an exact algebraic characterization of $D$-stability for dimensions $n > 4$ has remained a notoriously intractable open problem for over sixty years. In this paper, we establish a novel, systematic recursive framework that decomposes the structural check of $D$-stability into an analytical tree of parameter-dependent determinants. By applying a recursive delete/zero reduction strategy, we derive exact recurrence relations for the real and imaginary parts of the characteristic polynomial components. These algebraic relations uncover a structured hierarchy of new sufficient conditions for $D$-stability, expressed explicitly in terms of the matrix's principal minors. We show that while general numerical methods face unavoidable conservatism near the topological boundaries of the stable manifold, our deterministic framework provides sharp, absolute certification for low-order boundary matrices.

math.SP↗

Dynamical spin-nematic correlation in a transverse field Ising chain with non-Hermitian Gamma interaction

We investigate the effect of non-Hermitian Gamma interaction on the phase transitions and magnetic correlations for the transverse field Ising chain. We demonstrate that apart from the gapped antiferromagnetic and paramagnetic phases, there is a gapless phase induced by parity-time symmetry breaking, where the system exhibits long-range and short-range spin-nematic correlations in different regions divided by the quantum critical line determined from the correlation function and the subsystem entanglement entropy. Furthermore, we reveal that the parity-time symmetry breaking leads to the emergence of dynamical spin-nematic correlation, which also suggests a way of characterizing the spin-nematic map through non-equilibrium dynamics. Our findings show rich quantum phases stem from the competition among the Ising interaction, transverse field and non-Hermitian Gamma interaction, as well as providing a scheme for generating spin-nematic correlation in the spin chain.

cond-mat.mes-hall↗

A Census of Na D-traced neutral ISM and outflows at $0.6<z<4$

We present a statistical census of the Na D-traced neutral interstellar medium (ISM) and outflows in 309 galaxies at $0.6 10$), and 12\% in lower-mass systems. At high mass, ISM absorption is seen in both star-forming and quiescent galaxies, whereas in lower-mass systems it is observed only in star-forming galaxies. In massive quiescent galaxies, Na D detectability appears linked to star formation history: it is preferentially detected in older systems with larger 4000 Åbreaks, and younger, rapidly quenching galaxies with strong Balmer absorption H$δ_A$. We identify Na D {\it outflows} in 25 galaxies, revealing a possible dichotomy in driving mechanisms between star-forming and quiescent galaxies. In star-forming galaxies, outflow properties correlate with star-formation properties, consistent with a star-formation-driven origin. In quiescent galaxies, however, outflows are not associated with residual star formation and often require more energy than such star formation can provide. Together with the high AGN fraction among outflow-detected quiescent galaxies, this suggests that AGN dominate Na D-traced neutral outflows in cosmic noon quiescent systems. We further identify four quiescent galaxies with possible AGN fossil outflows, suggesting that AGN-driven outflows can persist beyond the active accretion phase and may help maintain quiescence.

astro-ph.GA↗

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

The Platonic Representation Hypothesis posits that neural networks trained on different modalities (e.g., text and images) converge toward a shared representation of reality. If true, this has significant implications for whether modality choice matters at all. In this paper, we show that the evidence for this claim is substantially weaker than subsequent work suggests. The mutual $k$-nearest-neighbor metric used on 1024 text-image pairs in the original study captures only coarse structure. To keep the alignment from collapsing as one scales up the data, $k$ has to grow proportionally, undercutting the argument for fine-grained representational convergence. The reported increase in alignment with language model strength saturates for recent models. Moreover, the one-to-one text-image pairing favors alignment, while alignment decreases with non-bijective data. We further find that image and text representations indeed share coarse semantic structure, but neither stronger language models nor richer captions yield fine-grained alignment. Thus, multimodal representations share coarse structure without evidence of convergence to a shared representation -- arguably, full representational convergence would require fine-grained alignment.

cs.CV↗

Verbal tics in frontier language models: A critical review of current releases, research evidence, and public discussion

Repeated praise, canned reassurance, familiar contrasts, and conspicuous vocabulary are recurring subjects in discussions of large language models. Their interpretation depends on context: a conventional phrase may be useful, while a fluent answer may reinforce a false belief. This critical review examines linguistic habits and sycophancy across eight developer families: OpenAI, Anthropic, Google DeepMind, xAI, ByteDance, Moonshot AI, DeepSeek, and Xiaomi. We verify current public offerings against official release and API documentation, with an evidence cutoff of 1 October 2026. We synthesize research on lexical overrepresentation, stylistic variation, social warmth, and agreement, alongside benchmark methods and dated English and Chinese public discussions. The research reviewed documents recurring linguistic patterns and agreement that distorts judgment; comparable measurements of the newest releases are sparse in the retrieved set. Current user reports include both complaints and improved writing, with experiences varying by task and prompting. We propose separate measures of recurrence, contextual appropriateness, and belief distortion, with precise service records and language-specific annotation. This framework makes claims about writing quality and conversational reliability testable as model services change.

cs.CL↗

Learning from the Near Future: Temporal Self-Distillation for RLVR

Reinforcement learning with verifiable rewards (RLVR) is a core post-training recipe for reasoning models, yet pure on-policy learning can be inefficient when useful trajectories are difficult to discover or exploration narrows. Existing self-guided approaches largely reuse capability already available to the current or earlier learner. We instead ask whether learning can also make use of capabilities that emerge later in training: can a model learn from its own future self? We introduce temporal self-distillation, in which a policy receives guidance from a stronger later checkpoint of itself. We hypothesize that the most useful temporal teacher need not be the strongest one: a teacher must provide sufficiently new capability while remaining compatible enough for that capability to be readily transferred, motivating a near-future regime. We study this principle through two complementary mechanisms. Near-Future Policy Optimization (NPO) performs off-policy behavioral transfer using verified future-self trajectories, while Near-Future Policy Distillation (NPD) performs on-policy token-level transfer on learner-generated trajectories. We further introduce AutoNPO, which adaptively determines when temporal guidance is useful and how far to roll back, turning future-self guidance into a repeated self-bootstrap process. Across eight image-text benchmarks, NPO improves GRPO from 60.25 to 62.84 and AutoNPO reaches 63.15, with consistent gains on text-only and video reasoning. Under NPD, a near-future teacher reaches 63.23 after continued RL versus 61.92 with a far-future teacher, despite lower immediate post-distillation performance. Together, these results suggest that effective temporal self-distillation depends not simply on teacher strength, but on a balance between newly acquired capability and learner compatibility.

cs.LG↗

SWE-chat: Coding Agent Interactions From Real Users in the Wild

AI coding agents are being adopted at scale, yet we lack empirical evidence on how people actually use them and how much of their output is useful in practice. We present SWE-chat, the first large-scale dataset of real coding agent sessions collected from open-source developers in the wild. The dataset currently contains almost 18,000 sessions, comprising more than 229,000 user prompts and 2 million agent tool calls. SWE-chat is a living dataset; our collection pipeline automatically and continually discovers and processes sessions from public repositories. Leveraging SWE-chat, we provide an initial empirical characterization of real-world coding agent usage and failure modes. We find that coding patterns are bimodal: in 41% of sessions, agents author virtually all committed code ("vibe coding"), while in 25%, humans write all code themselves. Despite rapidly improving capabilities, coding agents remain inefficient in natural settings. Only 59% of all agent-produced code survives into user commits, and agent-written code introduces more security vulnerabilities than code authored by humans. Furthermore, users push back against agent outputs - through corrections, failure reports, and interruptions - in 50% of all turns. By capturing complete interaction traces with human vs. agent code authorship attribution, SWE-chat provides an empirical foundation for moving beyond curated benchmarks towards an evidence-based understanding of how AI agents perform in real developer workflows.

cs.AI↗

Multicalibration for Unbiased Model-Based Prevalence Estimation

Estimating the prevalence of a category in a population using imperfect measurement devices (diagnostic tests, classifiers, or large language models) is fundamental to science, public health, and online trust and safety. Standard approaches correct for known device error rates but assume these rates remain stable across populations. We show this assumption fails under covariate shift and that multicalibration, which enforces calibration conditional on the input features rather than just on average, is sufficient for unbiased prevalence estimation under such shift. Standard calibration and quantification methods fail to provide this guarantee. Our work connects recent theoretical work on fairness to a longstanding measurement problem spanning nearly all academic disciplines. A simulation confirms that standard methods exhibit bias growing with shift magnitude, while a multicalibrated estimator maintains near-zero bias. While we focus the discussion mostly on LLMs, our theoretical results apply to any classification model. Two empirical applications -- estimating employment prevalence across U.S. states using the American Community Survey, and classifying political texts across four countries using an LLM -- demonstrate that multicalibration substantially reduces bias in practice, while highlighting that calibration data should cover the key feature dimensions along which target populations may differ.

cs.AI↗

Reionization, UV Luminosity and 21$\,$cm Sensitivity to Primordial Magnetic Fields: Impact of Energy Losses

Magnetic fields with field strengths between $10^{-17}\,$G and a few Nanogauss are expected to exist today in the intergalactic medium (IGM). Their origin is unknown, but may be of primordial nature, in which case they would have influenced the thermal and ionization history of the IGM as well as the growth of small-scale matter perturbations. In this work, we revisit constraints on Primordial Magnetic fields (PMFs) by consistently accounting for their energy losses through ambipolar diffusion and decaying turbulences from recombination through the epoch of reionization, which progressively reduces the magnetic field strength over time. We implement these effects in ${\tt HyRec}$ and ${\tt exo21cmFAST}$ to model the interplay between PMFs and astrophysical processes up to reionization. Using a neural-network emulator (${\tt NNERO}$), we perform a MCMC analysis that combines late-time probes of the reionization history and galaxy UV luminosity functions. We find that including PMF energy losses significantly relaxes previous bounds, as the reduced field strength suppresses their imprint on observables. Employing a Fisher matrix analysis, we estimate the sensitivity of the 21$\,$cm signal experiment HERA to the PMFs' imprint on intergalactic medium perturbations and show that 21$\,$cm cosmology could significantly improve on current bounds, depending on assumptions on the astrophysics. Our results highlight the importance of modeling PMF evolution self-consistently with the IGM evolution to extract current bounds and future sensitivities.

astro-ph.CO↗

The material life cycle of Koobor: separating coherent organization from potential-vorticity evolution in a wildfire-generated stratospheric vortex

Pyro-cumulonimbus convection associated with extreme wildfires can generate long-lived stratospheric vortices, yet their material coherence and life cycle remain poorly characterized. We apply geodesic vortex detection to winds from reanalysis to identify objectively defined (observer-independent) material boundaries of \emph{Koobor}, generated by the 2019--2020 Australian bushfires. These boundaries are closed material curves that undergo uniform finite-time tangential stretching and resist filamentation. We reconstruct their evolution across isentropic levels. Material coherence develops upward from 590~K, reaches maximum persistence at 690~K, and becomes progressively shorter-lived at higher levels. At 690~K, the reconstructed birth-to-death interval is approximately 40~days, while overlapping detections across levels define a nearly 60-day envelope. Analysis of potential vorticity (PV) within and around these material boundaries reveals a distinct dynamical evolution: the material boundary forms before strong PV isolation develops, a compact PV core emerges during the mature phase, and PV isolation weakens as material coherence approaches its demise. Strong PV isolation is concentrated at intermediate levels, whereas at 850~K material coherence occurs without a comparably well-defined PV core. These results provide an objective material description of the birth, vertical development, and decay of a wildfire-generated stratospheric vortex and show that material organization precedes the development of its strongest PV signature, which subsequently intensifies and weakens over the lifetime of the coherent vortex.

physics.ao-ph↗

Beyond average: heterogeneous first-passage dynamics in many-particle systems with resetting

Stochastic resetting is well understood for single-particle first-passage processes, but its consequences for collective first-passage behavior remain less clear. We address this problem in a many-particle system where all particles reset to the position of the rightmost particle, a protocol motivated by problems in artificial selection and avoidance. We use stochastic simulations of particles diffusing in a confining potential with an adsorbing boundary to examine two notions of group arrival: the first group-hitting time, when the first particle reaches the boundary, and the median group-hitting time, when half of them have reached it. We find that resetting produces broad hitting-time distributions with extended plateaus spanning several orders of magnitude. As the resetting rate increases, these plateaus extend and the mean group-hitting times grow rapidly. At the same time, the first-passage dynamics become increasingly heterogeneous. Consequently, the most probable and mean hitting times become widely separated, indicating the absence of a single characteristic time scale. These results demonstrate that the definition of group arrival is crucial for understanding and controlling collective first-passage behavior under resetting.

cond-mat.stat-mech↗

Fourier-Curve Constellations under Tangential Perturbation: Covariance-Aware Soft Demapping on Coded Links

A Fourier-curve constellation places $M$ points on a closed curve through $k$ complex slots. A Gaussian perturbation along the curve's tangent, whether injected as artificial noise or arising from first-order jitter of the curve parameter, gives every symbol an observation with a symbol-dependent rank-one covariance, and the maximum-likelihood symbol metric differs from the Euclidean rule by one rank-one correction per candidate. We realize this metric as a max-log soft demapper beside a Euclidean correlator bank at $2kM$ additional multiply--accumulate operations per symbol. On a regular $(3,6)$ LDPC-coded link at $(k,M){=}(20,64)$ it recovers $5.1$\,dB of the Euclidean mismatch at BLER${=}10^{-1}$ under natural labeling and $1.1$\,dB under Gray labeling, of which an average-covariance receiver recovers $0.7$\,dB and nothing measurable, respectively, and it makes the tangential perturbation $0.2$ to $1.0$\,dB cheaper than white noise of the same power on the same codebook. The per-tone phase orientation of the curve acts as an orthogonal rotation, so these results hold at every orientation, and the demapper decodes at the level of an exactly oriented receiver up to $0.2$\,rad of orientation error per component. A bit-interleaved coded-modulation achievable rate corroborates the ordering, a Woodbury extension keeps the rank-one structure under per-tone Rician fading, and $6$-bit lookup-table quantization costs no measurable degradation.

cs.IT↗

Domain-Adapted Small Language Models for Reliable Clinical Triage

Accurate and consistent Emergency Severity Index (ESI) assignment remains a persistent challenge in emergency departments, where highly variable free-text triage documentation contributes to mistriage and workflow inefficiencies. This study evaluates whether open-source small language models (SLMs) can serve as reliable, privacy-preserving decision-support tools for clinical triage. We systematically compared multiple SLMs across diverse prompting pipelines and found that clinical vignettes, concise summaries of triage narratives, yielded the most accurate predictions. The SLM, Qwen2.5-7B, demonstrated the strongest balance of accuracy, stability, and computational efficiency. Through large-scale domain adaptation using expert-curated and silver-standard pediatric triage data, fine-tuned Qwen2.5-7B models substantially reduced discordance and clinically significant errors, outperforming all baseline SLMs and advanced proprietary large language models (LLMs, e.g., GPT-4o). These findings highlight the feasibility of institution-specific SLMs for reliable, privacy-preserving ESI decision support and underscore the importance of targeted fine-tuning over more complex inference strategies.

cs.CL↗

Nanohertz gravitational waves from the baryon-dark matter coincidence

The nanohertz gravitational waves (GW) observed by pulsar timing arrays may originate from a cosmological first-order phase transition (PT) at $\sim$ 100 MeV. Taking this possibility seriously motivates the question: why 100 MeV? We point out that a PT at exactly those scales is predicted by the generation of the baryon asymmetry from a dark asymmetry via resonant neutron-dark matter oscillations, and we show that this PT can induce an observable GW signal compatibly with all experimental constraints. This proposal predicts dark matter self-interactions close to their observational upper limits and lowers the maximal expected mass of neutron stars. Independently of GW, this baryogenesis mechanism is tested by searches for missing-energy at the LHC and for neutron decays. We keep the model consistent with big-bang nucleosynthesis by adding heavy neutral leptons below 100 MeV, which generate neutrino masses and can induce further experimental tests.

hep-ph↗

Solving Hypergraph Laplacian Systems in Almost-Linear Time

For a connected weighted hypergraph, we give a randomized almost-linear-time solver for the Poisson problem for the cut-based hypergraph Laplacian in the natural input size $P=\sum_{e\in E}|e|$, the sum of hyperedge sizes. For every fixed constant $C>0$, our randomized algorithm runs in $P^{1+o(1)}$ time and, with high probability over its internal randomness, returns a primal point and a dual certificate, with additive optimality gap at most $\exp(-\log^C P)$. A key step is to rewrite the Fenchel dual as a convex-flow problem on an auxiliary $O(P)$-arc graph, yielding a near-optimal dual flow. The main difficulty is primal recovery, because this flow does not by itself determine a primal potential. Our main new ingredient is a recovery theorem showing that, for primal recovery, the detailed routing of the dual flow inside each hyperedge gadget can be discarded: one nonnegative scalar per hyperedge is enough. After the necessary finite-precision rounding, these scalars define a linear-cost min-cost-flow instance on the auxiliary graph, and solving it exactly recovers a primal potential. Finally, a ground-vertex reduction from regularized objectives to the Poisson solver gives randomized almost-linear-time resolvent/proximal primitives for the same cut-based hypergraph Laplacian.

cs.DS↗

Distributionally Robust Insurance under Bregman-Wasserstein Divergence

This paper investigates two optimal insurance contracting problems under distributional uncertainty from the perspective of a potential policyholder, utilizing a Bregman-Wasserstein (BW) ball to characterize the ambiguity set of loss distributions. The first problem examines an insurance demand model where the policyholder adopts an $α$-maxmin preference with Value-at-Risk (VaR). We derive the optimal indemnity function in closed form and study, both analytically and numerically, how the asymmetry inherent in BW divergence influences the optimal indemnity structure. The second problem employs a robust optimization framework, where the policyholder aims to secure robust insurance indemnity by minimizing the worst-case convex distortion risk measure while adhering to a guaranteed VaR constraint. In this context, we provide explicit characterizations of both the optimal indemnity and the worst-case distribution in closed form through a combined approach using the Lagrangian and relaxation methods. To illustrate the practical implications of our theoretical findings, we include a concrete example based on Tail Value-at-Risk (TVaR).

q-fin.RM↗