Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Copula Operad and Copula Entropy

We construct a symmetric operad $\mathfrak{C}$ on the class of all multivariate copulas, where operadic composition is given by Sklar substitution. We prove that the absolutely continuous subclass $\mathfrak{C}^{ac}$---which coincides with the $L^1$ class of copula densities---forms a suboperad; under composition, the density of the composite copula is given by the explicit Sklar substitution density formula $g(v)=ϕ\big(Ψ_1(v^{(1)}),\dots,Ψ_n(v^{(n)})\big)\prod_{k=1}^{n}ψ_{k}(v^{(k)})$. Furthermore, we show that copulas with finite copula entropy---identified with the $L\log L$ class of copula densities---are closed under substitution and hence constitute a suboperad $\mathfrak{C}^{L\log L}$. On this suboperad, copula entropy is strictly additive: $H(γ(Φ;Ψ_{1},\ldots,Ψ_{n}))=H(Φ)+\sum_{k=1}^{n}H(Ψ_{k})$.

math.PR↗

From Switching to Dynamic Regret: A Simple Reduction via Unbiased Random Sequences

In non-stationary online learning, dynamic regret has attracted increasing attention as a measure of how well an online learner performs against a time-varying comparator sequence. Despite considerable advances, attaining optimal bounds for strongly convex and exp-concave losses often involves intricate analysis. In this paper, we present a \textit{simple} framework that reduces dynamic regret minimization to switching regret minimization. As a result, we can derive dynamic regret bounds by using off-the-shelf algorithms with switching regret guarantees. The key idea of our reduction is to construct, for \textit{any} comparator sequence, an auxiliary random sequence that is unbiased at each round, with the controlled variance and a manageable number of switches. Combining this construction with suitable surrogate losses, we can decompose dynamic regret into the expected switching regret against the random sequence and its controlled variance. Theoretically, for strongly convex and exp-concave losses, we establish the $\widetilde{O}(T^{1/3}P_T^{2/3})$ dynamic regret bounds, where $T$ denotes the time horizon and $P_T$ denotes the path-length of the comparator sequence. Moreover, for general convex losses, the same reduction also recovers the $O(\sqrt{T(1+P_T)})$ dynamic regret bound. Notably, all our findings match the minimax optimal results for these three types of losses, highlighting the versatility of our proposed framework.

cs.LG↗

Chinese Competitive Debating Dataset and Benchmark

Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared rubric. We introduce a dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate at the match, stage, and speaker levels. We organized 182 matches and recruited 120 professional judges, with each match independently adjudicated by three judges using a predefined rubric. After excluding matches with incomplete records, the dataset contains 148 matches, 2,698 stages, and 20,542 exchange units, with manually verified transcripts and segmentation. It preserves original stage scores, match votes, best-debater ballots, and adjudication rationales. We define three tasks: winner-tendency prediction, stage-score prediction, and best-debater prediction. Zero-shot evaluation of multiple large language models yields a highest winner-prediction accuracy of 66.2%, a highest Pearson correlation of 0.250 between model stage scores and mean human ratings, and a highest best-debater prediction accuracy of 56.8%. The dataset and benchmark provide a testbed for studying large language models' understanding of interactive argumentation and their agreement with professional judges.

cs.CL↗

Exact Counts of Binary Phylogenetic Networks with Four Reticulations

Phylogenetic networks provide a flexible framework for representing reticulate evolutionary processes, such as hybridization, introgression, recombination, and horizontal gene transfer. However, their combinatorial complexity makes even basic enumeration problems difficult. Building on our previous work for networks with up to three reticulations, we derive an explicit closed-form formula for the number of unrestricted rooted binary phylogenetic networks with four reticulations on \(n\) labeled taxa. Our approach is based on tree-component graphs. We classify the 79 possible component graphs corresponding to networks with four reticulations into ten groups. We then enumerate the networks associated with each group by combining known counts of one-component networks, forests, and networks with fewer reticulations. Summing these contributions yields the desired formula. This result extends the exact enumeration of unrestricted binary phylogenetic networks to four reticulations and further demonstrates the effectiveness of component graphs for systematically organizing and counting increasingly complex network classes.

q-bio.PE↗

NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities

We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder and decoder-only language model with parallel specialized output streams for agent text and structured function calls, an auxiliary RNN-T branch for incremental user transcription, and a streaming TTS decoder. This design enables the model to listen, transcribe, reason, invoke tools, and speak within a unified streaming architecture while preserving the temporal behavior required for natural conversation. On Full-Duplex-Bench 1.0, NemotronLabs VoiceChat achieves the lowest pause-handling takeover rates among evaluated open-weight systems, 100\% takeover following user interruptions, and a 4.33/5 post-interruption response-quality score. On Full-Duplex-Bench 1.5, it resumes its response after user backchannels in 93\% of cases. NemotronLabs VoiceChat obtains a 55.1 normalized average on VoiceBench and, on Full-Duplex-Bench 3.0 (FDB 3.0), achieves 82.5\% tool-selection F1, while argument accuracy and end-to-end tool execution remain areas for improvement. These results demonstrate that full-duplex interaction, speech recognition and generation, general language capabilities, and external tool use can be integrated in a single open speech-to-speech model without sacrificing real-time conversational behavior.

cs.CL↗

A Channel-Boosted Multi-Agent System with Iterative Consultation for Document Sensitivity Classification

Organizations in critical national infrastructure sectors must assess heterogeneous documents for sensitivity before routing or storage. Manual assessment is slow, inconsistent, and unscalable. Extending our prior leakage-controlled benchmark, BERT established the top single-encoder baseline (89.14% accuracy, 89.33% F1-score under 5-fold cross-validation on the Strategic 16K corpus). However, transformer baselines suffer from a structural limitation: fixed input length truncation discards evidence beyond the retained window-precisely where sensitive cables tend to be longest. We present Channel-Boosted MAS (CB-MAS) and instantiate it as IC-MAS (Iterative Consultation Multi-Agent System) to solve this without long-context computational costs. A Channel Critic Agent learns document-adaptive trust weights governing Gated Channel Boosting between two first-window encoders, while paired Consultation Agents iteratively exchange belief states to reconcile evidence from the beginning and end of long documents. IC-MAS holds computation constant regardless of document length by reconciling fixed windows in a compact representation space. Ablation studies show critic-controlled Channel Boosting provides the bulk of accuracy gains, while consultation recovers recall without precision collapse. Critic-Controlled Gated Channel Boosting with Max-Pool fusion and Blackboard Adaptive Consultation achieves 90.72% accuracy, 91.23% F1-score, 92.01% sensitive recall, and 90.46% sensitive precision, using about 54% less average computation than a fixed-round baseline. Gains over the single-encoder baseline are statistically significant (McNemar's test, p less than 0.000001; paired t-test). We include LIME/SHAP explainability, multi-agent evaluation, and an honest accounting of limitations.

cs.CL↗

Microcanonical Operator Spectrum, BPS Chaos and Free Probability Theory

Matrix elements of simple operators are central to thermalization in chaotic quantum many-body systems. Their moments are often described statistically through the Eigenstate Thermalization Hypothesis. In this paper, we study a more refined characterization of these matrix elements: the eigenspectrum of simple operators once projected to microcanonical energy windows. The spectrum encodes all moments of the matrix elements at once, and is particularly relevant when the spectrum of the Hamiltonian is degenerate. Assuming that the eigenbasis of the simple operator is Haar-randomly oriented with respect to that of the Hamiltonian, we prove that the spectrum of an operator projected to a small microcanonical window obeys Gaussian random matrix statistics. We do this directly by studying the probability distribution inherited from the random orientation, as well as using tools from Free Probability Theory. We discuss the implications for BPS chaos and more generally for eigenstate thermalization.

hep-th↗

Extremal Least Common Multiples in Rows of Pascal's Triangle

For $r \ge 0$, let $\mathcal P_r=\{\binom{r}{0},\binom{r}{1},\ldots,\binom{r}{\lfloor r/2\rfloor}\}$ be the set of distinct entries in row $r$ of Pascal's triangle. We study the least possible least common multiple of $n$ entries chosen from one row, with the row itself also free: $a(n)=\min_{r\ge0,\ S\subseteq\mathcal P_r,\ |S|=n}\operatorname{lcm}(S)$. Whereas the least common multiple of an entire row is given by a classical identity of Farhi, allowing both the row and the selected coefficients to vary creates a different optimization problem. We first recast the fixed-row problem exactly as a weighted prime-power exclusion problem, which explains how omitting a few coefficients can remove expensive prime-power contributions. Our main asymptotic result is $\log a(n)=2n+O\!\left(n\exp\!\left(-c(\log n)^{3/5}(\log\log n)^{-1/5}\right)\right)$ for some absolute $c>0$, so $a(n)^{1/n}\to e^2$. A two-band refinement further shows that every optimal row satisfies $r_n=2n+O\!\left(n\exp\!\left(-c(\log n)^{3/5}(\log\log n)^{-1/5}\right)\right)$, and that the minimum prefix defect of an optimal support is $o(n)$.

math.CO↗

Global Existence and Large-Time Asymptotic Behavior of Strong Solutions to the Two-Dimensional Nematic Liquid Crystal Flows with Large Initial Data and Vacuum

This paper studies the two-dimensional nonhomogeneous incompressible nematic liquid crystal flows with planar orientation fields taking values in $\mathbb{S}^1$. For the initial-boundary value problem in bounded domains with density-dependent viscosity, we establish the global existence and exponential decay of strong solutions with initial density allowing vacuum, provided that $\|\nabla μ(ρ_0)\|_{L^q}$ is sufficiently small for some $q>2$. The smallness condition is automatically satisfied when $μ$ is constant, and hence the result yields the global existence of strong solutions for arbitrarily large initial data in the constant viscosity case. Furthermore, for the Cauchy problem with constant viscosity and either vacuum or non-vacuum far-field density, we establish the global existence and large-time decay rates of strong solutions for arbitrarily large initial data. These results are obtained without imposing any geometric angle conditions on the initial orientation field.

math.AP↗

SyzHarness: Patch-Based Kernel Bug Reproduction with LLM-Synthesized Fuzzing Harnesses

Automated kernel vulnerability reproduction is essential for bug triage, patch validation, and regression testing, but still lacks an effective and efficient solution. The core challenge is twofold: a reproducer must first recover the trigger scaffold needed to reach the vulnerable state and determine the precise concrete values that actually trigger the bug. Existing directed fuzzing approaches are ineffective at recovering the necessary trigger scaffold, while LLM-only generation is brittle because it struggles with concrete-value discovery and runtime nondeterminism. We design SyzHarness, a framework that combines LLM reasoning with coverage-guided fuzzing for patch-based Linux kernel vulnerability reproduction. Given a patch, SyzHarness uses an LLM agent grounded by code navigation tools to synthesize a parameterized fuzzing harness that fixes the prerequisite setup logic while exposing only uncertain, bug-critical input parameters to be mutated by Syzkaller. SyzHarness then translates this harness into a Syzkaller compatible interface and iteratively refines it using hierarchical reachability feedback. We evaluate SyzHarness on multiple datasets of triggerable real-world Linux kernel vulnerabilities. On 100 KernelCTF cases, SyzHarness achieves a 78% bug reproduction success rate. On the SyzDirect benchmark, SyzHarness achieves a 73% bug reproduction success rate, substantially outperforming prior directed greybox fuzzing. On 50 recent, known-triggerable syzbot bugs fixed after March 2026, SyzHarness reproduces 40/50 (80%) using only the fix commits as input.

cs.CR↗

LumoTree: Path-Parallel Speculative Verification for Hybrid Language Models

Tree speculative decoding for hybrid language models must preserve one coherent continuation across recurrent, convolution, and attention state. We present LumoTree, a verifier that executes recurrent paths in parallel, reuses state tiles within each path, and coordinates native recurrent replay, convolution-history gathering, and attention-cache remapping through a shared logical tree. Fused candidate selection, GPU-resident acceptance, and grouped split-K attention support the verification cycle. Component experiments show exact candidate-selection parity and recurrent agreement within paired error bounds. An exploratory Qwen3.8-27B NVFP4 deployment on a single NVIDIA DGX Spark (GB10) records 25.63 pooled tokens/s on ten SWE-bench Verified Astropy tasks. The results characterize component-level numerical agreement and coding-agent deployment; complete-model continuation and controlled application speedups remain open.

cs.LG↗

Repeated differentiation of random polynomials with i.i.d. rotationally invariant roots

Let $p_n$ be a random polynomial of degree $n$ whose roots are independent and identically distributed according to a rotationally invariant probability measure $μ_0$ on the complex plane with finite logarithmic moment. If $k_n/n\to t\in(0,1)$, we prove that, as $n \to \infty$, the empirical zero measure of the $k_n$-th derivative of $p_n$ converges weakly in probability to a deterministic rotationally invariant probability measure. We describe the limiting measure explicitly in terms of the radial quantile function of $μ_0$. This proves a conjecture of Hoskins and Kabluchko [Exp. Math. 32 (2023), no. 4].

math.PR↗

Video-STLayout Pre-training

In recent years, pre-training has become fundamental to learning effective video representations, enabling strong transfer to downstream tasks. A popular framework in pre-training involves aligning features of a video encoder with that of another modality, for example, language or audio. We introduce Video-STLayout pre-training, a novel strategy for obtaining rich video representations informed by spatio-temporal layout of object bounding boxes. Object layouts can easily be obtained by applying an off-the-shelf object detector on the video frames. Our method uses a contrastive loss to align video features with the layout features from a trained layout encoder. We show the effectiveness of our approach in the task of activity recognition in complex scenes.

cs.CV↗

An Equivalence of Categories between 3-Crossed Modules and Gray 4-Groups

In this paper, we investigate the relation between the category of 3-crossed modules and the category of Gray-type 4-groups. The notion of a 3-crossed module was first introduced by Arvasi \textit{et al.}, motivated by the question of what kind of algebraic structure completely encodes a homotopy 4-type. On the other hand, from the point of view that higher groups are equivalent to algebraic realizations of higher categories -- as exemplified by the relationship between 2-crossed modules and Gray 3-groups established by Sarikaya--Ulualan -- it had not been clear how the 3-crossed modules of Arvasi \textit{et al.} relate to any higher category. In our previous paper, we proposed a new definition of a 3-crossed module and observed that it admits a natural interpretation in terms of higher categories. In this paper, we make this interpretation precise: we introduce a 4-category, which reduces to a semistrict braided monoidal 2-category when restricted to a single object and a single 1-morphism, and prove that the category of our 3-crossed modules is equivalent to the category of Gray 4-groups, defined as single-object versions of this 4-category in which all morphisms are invertible. Furthermore, while Gray 4-groups provide a conceptual framework for 4-dimensional higher structures, they are often computationally intractable. Our equivalence establishes 3-crossed modules as a concrete, group-theoretic calculus for Gray 4-groups, providing a powerful computational tool for studying surface knots, state-sum invariants, and higher gauge theories.

math.CT↗

Unistochastic reduction of multi-matrix coherent-state kernels

We study finite-$N$ coherent-state overlap kernels in multi-matrix BPS sectors of $\mathcal N=4$ super Yang--Mills theory. For commuting coherent-state parameters, the kernel is a Laplace transform of the Haar-pushforward measure on the unistochastic set. After quotienting row- and column-shift redundancies, the nontrivial source is an $(N-1)\times(N-1)$ matrix $D$. On the physical $m$-matrix locus, $D$ is a Gram matrix with $\operatorname{rank}D\leq m$, and the ordinary HCIZ sector is precisely the rank-at-most-one stratum. For $N=3$ we obtain an exact Bessel-integral representation of the first genuinely rank-two kernel. Along the rank-one ray $D=rE_{11}$, $F_N={}_1F_1(1;N;r)$ for arbitrary finite $N$. Using $K_N=\log F_N$ as the reduced Kähler potential, we study a controlled two-matrix embedding of the HCIZ sector and determine its transverse metric, second fundamental form, and Gauss--Codazzi geometry. The embedding is non-totally-geodesic at finite $N$. In the scaling $r=Nρ$, we find \begin{equation} Nh_N\to\min(ρ,1), \qquad \|B\|^2\to2(1-ρ)_+. \end{equation}

hep-th↗

On Emergent Capabilities and Model Merging

Fine-tuned checkpoints and adapters now fill public repositories, and the most common operation applied to these artifacts is model merging: arithmetic on their weights that assembles capabilities cheaply. We ask what this operation does to emergent capabilities: behaviors an artifact carries that were never an explicit training target. Studying two independent testbeds (activation oracles and emergent-misaligned models) across three model families, we find that the answer is threefold. First, merging preserves an emergent capability that both parents carry: merging two misaligned checkpoints retains most of their broad misalignment across the whole mixing range. Second, merging cannot create an emergent capability that is superadditive in its parents: no weighted merge of two single-task oracles reaches the jointly-trained oracle's auditing ability. Third, when only one parent carries the capability, merging dilutes it faster than the trained capability that accompanies it: the gap is significant in most settings. In short, emergent behaviors of an artifact do not compose the way its trained capability does.

cs.LG↗

Constraints on admissible behavior of GMRES applied to tridiagonal Toeplitz systems

The result of Greenbaum, Pták, and Strakoš that for a given set of eigenvalues, any convergence curve is possible [SIMAX 1996] and the subsequent parameterization of such matrix-right-hand side pairs $(A,\mathbf{b})$ of Arioli, Pták, and Strakoš [BIT 1998] demonstrated that the behavior of the \gmres could not be completely characterized by the eigenvalues of $A$ alone. In this paper, we consider how to use this theory to understand the admissible and attainable \gmres behavior for matrices with constrained structure, focussing on non-Hermitian (non-symmetric) tridiagonal Toeplitz matrices. We show that Toeptliz structure necessarily constrains the how the theory from these papers can manifest but that a continuum of admissible behaviors is still attainable.

math.NA↗