Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Likelihood Ranking doesn't Scale Like Prompting in LLMs

LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer set, and can therefore leverage the same task-conditioned answer-selection interface used in prompting. We study a complementary protocol based on likelihood ranking of declarative statements constructed from the same question--answer pairs. Across 95 decoder-only models, ranging from 0.1B to 104B parameters, and 10 MCQA datasets, we find a systematic divergence between declarative-statement likelihood ranking and prompted answering. Statement-likelihood accuracy remains comparatively stable across scale, whereas prompted answering improves sharply with scale and instruction-tuning. These results suggest that likelihood preferences over controlled declarative alternatives and task-conditioned answer selection probe distinct aspects of model behavior, and should not be treated as interchangeable.

cs.CL↗

Rufus-Air: An Open LLM Post-Training Recipe

Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe. Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals. Training builds on open-source components and public data, much of it used as released, without new human annotation or an in-house distillation teacher. Our main findings are that (i) diverse, high-quality SFT establishes a strong capability floor; (ii) difficulty filtering keeps RL prompts within a productive learning range; (iii) reward reliability provides a practical principle for ordering stages; and (iv) infrastructure and engineering choices are part of the recipe, not just an implementation detail. Rufus-Air improves over the official GLM-4.5-Air post-trained release and is competitive with similarly sized open models.

cs.CL↗

Detecting Glaucoma Across Multi-ethnic Myopic and Non-Myopic Populations Using an Uncertainty-Aware Vision Transformer: A Multicentre Model Development and Validation Study

Background: Artificial intelligence (AI)-based glaucoma detection from colour fundus photographs (CFP) offers scalable screening, but performance may decline on external datasets because of differences in ground-truth definitions, populations, and coexisting conditions such as high myopia (HM). We developed and validated a Vision Transformer-based deep learning (DL) model for glaucoma detection across multi-ethnic cohorts with and without HM. Methods: A ViT-B/16 model with predictive uncertainty estimation was developed using 56,483 CFPs (57.1% with myopia; 14.4% with HM). Glaucoma labels were standardised using clinical, imaging, and perimetry data. The model was validated on 16 independent datasets across three continents, including four datasets with explicit HM labels. Findings: Internal AUROC was 98.7% (95% CI 98.2-99.1%), with sensitivity 94.5% and specificity 97.3%. Across 16 external datasets from eight countries, AUROCs ranged from 86.4% to 99.6%. In HM eyes, internal AUROC was 97.8% (95% CI 96.1-99.2%), with sensitivity 94.8% and specificity 93.7%. External HM AUROCs were 86.5% in the Beijing Eye Study and 93.3%, 91.8%, and 85.5% in hospital-based datasets from Taiwan, Thailand, and South Korea. In an exploratory HM clinical evaluation, the model had higher CFP-only diagnostic accuracy than ophthalmologists and trained graders (92.0% vs 70.0%; p=0.008) and performed comparably to glaucoma specialists using full clinical information. Interpretation: The model showed robust glaucoma detection across myopic and non-myopic multi-ethnic populations and may support AI-assisted screening in settings with high HM prevalence.

cs.CV↗

ToCo-Mesh: Topology-Consistent Dynamic Mesh Reconstruction via Adaptive Tessellation and Surface-Aligned 2DGS

Reconstructing dynamic meshes with consistent topology from multi-view temporal images remains a challenge. Existing approaches typically face a dilemma between fine-scale shape recovery and topological stability. Frame-by-frame extraction methods capture fine details but break vertex correspondence, leading to flickering meshes. Conversely, template-based deformation ensures consistency but struggles to adapt its surface resolution during optimization, missing local surface details. To address these limitations, we propose ToCo-Mesh, a dynamic reconstruction framework that maintains topology consistency over time while achieving high-fidelity geometry. Specifically, we introduce a dual-mesh representation, where a canonical template mesh is tightly bound to time-varying coarse guide meshes via barycentric parameterization. While keeping guide meshes fixed to condition the deformation, we perform error-driven split-and-merge on the template mesh to progressively increase reconstruction fidelity. Furthermore, to suppress surface irregularities and achieve photorealistic rendering, we incorporate a Surface-Aligned 2DGS module. By anchoring flattened Gaussians to mesh faces, we utilize their rendered normals to guide inverse geometric fine-tuning. To our knowledge, ToCo-Mesh is the first framework to enable adaptive mesh refinement while maintaining strict topological consistency. Extensive experiments demonstrate that our method achieves SOTA geometric accuracy while maintaining competitive rendering quality.

cs.GR↗

A Periodic Long-Time Boltzmann--Grad Limit in Every Fixed Dimension $d\ge 4$

We prove a periodic long-time Boltzmann--Grad limit for hard spheres in every fixed spatial dimension $d\ge4$. Under the assumptions of the main theorem, the rescaled $s$-particle correlations converge in $L^1$ to the corresponding tensor-product Boltzmann profile at rate $\varepsilon^{1/(400d)}$, uniformly for $1\le s\le|\log\varepsilon|$ and $0\le t\le t_{\rm fin}$. In particular, when the activity and weighted solution bounds are fixed, $t_{\rm fin}=O(\log|\log\varepsilon|)$. The componentwise long-bond estimate used in dimensions two and three is insufficient in higher dimension. We replace it by a joint estimate for two connected time sublayers, selecting the first two lower collisions as landing roots. Disjoint landing edges yield a direct tangential frame. When the edges overlap, the newly appearing particle line contains a collision-free chord whose radial relative speed produces a rank-$(d-1)$ positive Schur factor. A positive Jacobi-network argument prevents focusing, while bounded-degree finite-fibre coarea globalizes the estimate for arbitrary nonnegative joint kernels. A first-failure construction repairs periodic double overlaps, and the resulting packet bound closes the exceptional top-layer contribution. A paid epoch restart for sealed complete joint kernels, using the fixed-word multi-landing operator for every fixed integer $k\ge1$, yields the stated time range. Within the regular connected full two-sublayer packet class, every packet with at least $2k-1$ physical lines admits a canonical $k$-birth flag and the operator factor $\varepsilon^{(d-1)(k-1)}\varepsilon_*^{-k(d-1)}$; $k=2$ is the first gain-producing case. A complementary full-chord estimate provides a coarse fallback.

math.AP↗

Two questions on $E_0$-like generic equivalence

We answer Problems 8.1 and 8.3 from \cite{Tianyuan2026}. We first note that the condition attributed there to ordinary Prikry forcing is not correct as written: the appropriate formulation uses finite symmetric difference of the ranges of the generic sequences, rather than eventual equality at the same coordinates. We record answers to the two problems for both formulations. A length-$ω$ Magidor forcing gives a genuinely Prikry-type example satisfying both conditions. Rigid real-adding forcing gives a broad source of further examples, and a forcing of Jech and Shelah shows that the condition does not carry any large cardinal strength. An $E_0$-invariant Jensen-type forcing of Kanovei and Lyubetsky gives a stronger nontrivial example in which the generic reals in a fixed extension form exactly one full $E_0$-class. Finally, a simple recoding turns the real-forcing previous examples into answers for the version of the problems with corrected condition.

math.LO↗

Beyond Harnack Rigidity

The boundary of a plane amoeba is always contained in its contour, and equality is a characteristic feature of simple Harnack curves. We show that the converse fails, even under strong smoothness and nondegeneracy assumptions. For every two-dimensional lattice polygon, except unimodular triangles, we construct a smooth Newton-nondegenerate curve with smooth logarithmic critical locus and smooth embedded contour satisfying $\mathcal C(\mathscr A_f)=\partial\mathscr A_f$, although the curve is not Harnack. We also provide explicit primitive and nonprimitive families that are not torus-equivalent to simple Harnack curves. A concrete primitive example is certified by exact elimination and Sturm root counting. These results disprove contour--boundary rigidity and show that the contour as a set does not detect the real structure or the multiplicity of coincident critical sheets, thereby refining the compensation problem proposed by Lang, Shapiro, and Shustin.

math.AG↗

Dimer model and random lattice permutations with general weights

The dimer model and random lattice permutations are two fundamental objects at the interface of probability, combinatorics, and mathematical physics. We study these models on finite periodic boxes in $\mathbb Z^d$ within a common framework. For the dimer model, edges connecting arbitrary vertices carry a weight which depends on their relative displacement and dimer configurations are weighted through their occupied edges. Superimposing two independent perfect matchings gives the double-dimer model, whose configurations are collections of disjoint loops. Permutations, instead, are weighted through the spatial displacement of their jumps. For broad classes of weights of finite or infinite range we prove long-range order and the occurrence of macroscopic loops. This extends nearest-neighbour results to arbitrary-range edges and jumps. In particular, long-range weights yield long-range order and macroscopic loops already in dimensions $d=1,2$. In dimension two, this behaviour is qualitatively different from that of the nearest-neighbour model. We complement these results with sharp absence criteria.

math.PR↗

DAMSEP: Distance-Aware Monaural Source Separation using Multi-RIR Estimation

Although room impulse responses (RIRs) encode source-distance cues, conventional monaural source separation focuses on recovering audio content without estimating source-specific RIRs, losing the associated spatial information. To address this limitation, we propose Distance-Aware Monaural Source Separation using Multi-RIR Estimation (DAMSEP), the first end-to-end framework that is jointly trained for source separation and multi-source RIR estimation from a single-microphone mixture. DAMSEP integrates a separation backbone with shared dereverberation and RIR estimation modules to jointly recover clean sources and source-specific complex convolutive transfer functions under source estimation and reverberant reconstruction objectives, enabling relative near/far ordering through the direct-to-reverberant ratios of the corresponding RIRs. For comprehensive evaluation, we introduce HETMIXR, which spans heterogeneous source content and diverse simulated room conditions with source-specific RIRs and geometric distance annotations. Experiments on HETMIXR demonstrate superior performance in source separation, RIR estimation, and distance ordering. Ablation studies reveal the complementary benefits of source supervision and reverberant reconstruction, while additional evaluations show generalization to single-speaker inputs and mixtures generated using measured RIRs from an unseen room. Our code and dataset are available at https://github.com/Wenanzhi/DAMSEP.

eess.AS↗

Joint Analysis of Latent Dimensionality and Frame Rate in Continuous Audio Encoders

Continuous audio encoders compress audio along feature and time axes through latent width and frame rate, but their joint effect on downstream performance remains unclear. We train sixteen encoders spanning four widths and four frame rates, with downstream adapters and probes, using matched training protocols. Despite generally improved reconstruction at larger widths, automatic speech recognition (ASR) and spoken question answering (SQA) favor moderate widths at higher rates, with the best observed widths shifting toward larger values under stronger temporal compression. Frozen-model PCA interventions reveal distinct reconstruction and recognition sensitivities: removing the trailing half of the components substantially degrades ASR in high-rate 512-dimensional encoders with comparatively small reconstruction penalties, whereas 1024-dimensional encoders largely preserve both. Yet the projected 1024-dimensional model underperforms unmodified narrower models on ASR at 12.5Hz. These findings identify a width--rate interaction in downstream utility and suggest that how representations are organized during training matters beyond reconstruction fidelity and compressibility.

cs.SD↗

PUBG Ally: A Conversational Embodied Agent as an AI Teammate

We introduce PUBG Ally, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate. Building such a teammate requires combining two difficult capabilities: it must perceive and respond to a constantly changing game world under strict latency constraints while interacting naturally with players, keeping its speech synchronized with its actions. Ally therefore combines agentic tool use with real-time game control. A language-model agent uses a controlled interface to inspect game information, interpret player speech, maintain context, decide what to say, and issue high-level action choices that steer a faster control layer for movement, combat, and recovery. Because the player's and Ally's speech and actions continually shape each other and the course of the match, training requires data from actual gameplay. We therefore collect data across nearly 39k sessions in which real players play alongside Ally, recording gameplay, player speech, agent decisions, tool use, actions, and player feedback, and use these records for iterative training. To evaluate teammate quality, we use player feedback and preference comparisons to identify gaps between offline evaluations and player preferences, and iteratively refine the evaluation criteria. Deploying Ally in live service further requires low-latency on-device execution and safeguards for player-facing communication, which we address through model compression, context compaction, targeted safety training, runtime guardrails, and memory redaction. During the live service, we surveyed players in 141 countries. Among respondents whose play with Ally was confirmed in game records, positive responses exceeded negative responses by 25.1 percentage points when asked whether they would recommend Ally, with players describing Ally not only as a tool but also as a teammate or companion.

cs.AI↗

The Two-Component Becker-Döring System: Basic Properties

The two-component Becker-Döring model is a two-dimensional coagulation-fragmentation equation that describes the dynamics of clusters built from two different types of monomers growing and shrinking through monomer interaction. We show the existence and mass conservation of solutions and give a partial result regarding uniqueness. Furthermore, we prove under the assumption of detailed balance the entropy equation, followed by a detailed study of the equilibrium states and relative entropies. We can determine the minimum of all relative entropies under sequences of fixed Type I and Type II masses and fully describe the minimising sequences in the weak* topology.

math.AP↗

Long Time Behaviour of the Two-Component Becker-Döring System

The classical Becker-Döring equations describe the formation of clusters by aggregation and fragmentation of monomers. If the total amount of mass is supercritical, larger and larger clusters are formed, leading to an asymptotic loss of mass in the long time limit. The two-component Becker-Döring system arises when clusters are built from two types of monomers. Here, we study the natural extension of the one-component system, where no energy or entropy is leaving or entering the system - the so called detailed balance assumption. We rigorously prove that all initial conditions admit a solution minimising the relative entropy as time approaches infinity. This shows, that the long time limit in the weak* topology is selected through the initial Type I and Type II masses. Furthermore, the proportion of lost Type I mass is determined by the limit point via the mixing ratio of Type I and Type II monomers that maximises the binding energy. The main difficulty of the two-component system is that no relative entropy is weak* continuous, which is crucial for the classical argument [2]. Instead, our proof is based on a discrete two-dimensional logarithmic Sobolev inequality, that bounds the entropy dissipation on the correct timescale. Our approach also improves current assumptions to determine the long time behaviour for the one-component system.

math.AP↗

A Computer-Assisted Proof of Speed Monotonicity for the Biased Random Walk on a Galton-Watson Tree Beyond the Known Range

The speed v(lambda) of the lambda-biased random walk on a supercritical Galton-Watson tree without leaves is conjectured to be nonincreasing on [0,m), where m is the mean offspring. For non-constant offspring, monotonicity is known only for small bias: lambda <= 1/1160, lambda <= 1/2, and, when every vertex has at least m_1 >= 2 children, lambda <= m_1/(1+sqrt(1-1/m_1)). For offspring uniform on {2,3} (m=2.5) the last bound is 2/(1+sqrt(1/2)) = 1.17157... We prove, with computer assistance, that v is strictly decreasing on [0,1.755] for this law. The proof has three parts. Aidekon's speed formula gives v=(R-lambda)/(R+lambda) for an explicit functional R, so v decreases exactly when R/lambda does; we compare R/lambda at two biases directly, which avoids differentiating the conductance. A pathwise Lipschitz bound on the conductance in lambda turns that comparison into an inequality between expectations of explicit functions. A monotone sandwich of discretised laws gives two-sided bounds on the conductance law, and each lambda-cell is verified with exact rational arithmetic on top of bounded floating-point error; an independent interval-arithmetic implementation re-certifies five cells, including the one that sets the endpoint. The method stops where the crude Lipschitz bound becomes too weak; sharper control of the derivative of the conductance is what the full range needs. Code and certificates are public.

math.PR↗

Aim Short to Reach Far: Your Frozen World Model Can Plan Better Than You Think

Planners built on visual world models commonly score each predicted outcome by its distance to the encoded goal image. We show that this target can limit control even with exact dynamics and globally optimal short-horizon search: reaching a goal may require actions that initially move away from it. With frozen LeWM models, intermediate targets substantially improve action synthesis and recorded-action ranking on Cube, PushT, Reacher, and TwoRoom. Learned targets and targets drawn from observed experience both produce these gains. We introduce Anchored Planning, which retrieves a recorded segment whose start and end resemble the current and goal observations, then aims at an observation shortly after its start. The frozen model scores actions toward this target from the current state. Without additional training, planning toward observed targets outperforms the LeWM planner on every task in our long-range evaluation. Additional final-goal search falls short of the same gains. Lower successor-prediction error need not translate into better control. Success also depends on how far ahead the target is placed and on shrinking the retrieval span as execution advances. Changing only the target lets the same frozen model and planner reach goals that final-goal scoring misses.

cs.LG↗

PK/PD-integrated Bayesian platform design for phase II dose regimen optimization

Early-phase dose-finding methods increasingly assess toxicity and efficacy jointly, but comparisons based only on administered dose may inadequately characterize regimens differing in schedule. We developed a Bayesian phase II adaptive platform design for regimen optimization that integrates pharmacokinetic/pharmacodynamic (PK/PD) modelling into toxicity, efficacy, regimen selection and adaptation decisions. The proposed PK/PD-informed Regimen Optimization Platform (PROP) design uses a population PK/PD model to generate patient- and population-level predictions of exposure and biological activity. Acute and cumulative toxicities are analysed using a discrete-time time-to-event model informed by PK exposure. Efficacy is evaluated through Bayesian model averaging of exposure-driven and biomarker-driven time-to-event models. The design supports regimen graduation, discontinuation for futility or safety, and addition of unexplored regimens. Performance was evaluated through simulations motivated by an influenza intensive-care setting. Across six scenarios, PROP generally improved graduation and futility decisions, reduced inappropriate graduation, and supported the addition of promising regimens compared with dose-based alternatives. It also more accurately estimated regimen-specific toxicity and arm-specific efficacy, while the model-averaging framework favored the efficacy model consistent with the data-generating mechanism. Dose-based approaches performed better for safety stopping in some scenarios, despite less accurate characterization of the regimen--toxicity relationship. PK/PD-informed platform designs can improve adaptive regimen selection and knowledge generation when dose alone cannot adequately characterize treatment regimens.

stat.ME↗

PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations

Large language models increasingly operate as persistent assistants in user-facing, shared-session, and tool-augmented settings. When users disclose sensitive information during an active conversation, that information may remain behaviorally recoverable through later prompts even after the dialogue shifts to unrelated topics. We introduce PrivDrift, a benchmark for auditing whether user-disclosed secrets remain recoverable after conversational topic drift and persuasion-based probing. PrivDrift contains 1,000 controlled multi-turn dialogues with seeded secrets, content-dense drift turns, and standardized extraction probes. Across three LLMs with extended context windows, dialogue-level hybrid leakage remains substantial, ranging from 38.7% to 54.6%, and varies strongly by model, secret type, and persuasion intensity. Within the tested drift window, additional topic drift does not reliably reduce leakage, suggesting that privacy risk in active LLM contexts should be evaluated as a persistent behavioral failure mode rather than only as training-data memorization or immediate jailbreak behavior.

cs.AI↗