Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 595 records · Page 33Linked to original sources

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

Scientific discovery begins where known problems end. There, AI systems must engage in exploration: framing hypotheses, designing experiments, and iterating on the results. However, evaluating this ability is difficult: (1) how to verify whether a genuinely new hypothesis holds, and (2) how to determine whether a system has discovered it through exploration or merely recalled related knowledge from pre-training data. To this end, we introduce ExplorationBench, which turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds: their rules are executable, so every answer can be checked exactly, and they conflict with familiar knowledge, so recall alone cannot solve the tasks. The benchmark contains two sandboxes, AlienCode (31 discovery targets, 70 tasks) and AlienLogic (24 discovery targets, 70 tasks). Each sandbox provides a flawed manual, task-specific environmental feedback, and a dedicated tool-call schema. Systems use these resources to explore the sandbox, then solve held-out tasks. We evaluate 10 AI systems and find that the strongest systems can acquire and apply unfamiliar rules, while performance varies substantially across trajectories and continued exploration can stall or reverse earlier gains. ExplorationBench represents a step towards AI systems that can acquire and apply genuinely new knowledge through exploration in unknown environments.

cs.AI↗

JevOut: Natural Context Can Flip Decision Models

An ordinary-looking background detail can turn a correct model decision into a confident mistake. We demonstrate this fragility in four decision systems, including Jev, across seven datasets covering knowledge, reasoning, and tool routing. Within 64 accepted target evaluations per decision, we uncover short context additions that redirect 61.4%-73.2% of each system's initially correct decisions toward a wrong option fixed in advance. The additions supply background or procedural information rather than explicit answer-selection instructions, leaving the original question and choices intact. We construct them through probability-guided context optimization, which uses shifts in the option distribution to refine surrounding text under naturalness and answer-preservation constraints. Redirection affects initially confident decisions, often produces high-confidence wrong choices, and transfers across models. In blinded human evaluation, 91.6% of 250 sampled successful contexts are judged natural, answer-preserving, and free of decisive answer-changing evidence by a majority of three independent annotators. These findings expose a weakness in current decision models: context that looks entirely compatible with an input can redirect the choices that agents, routers, and evaluators rely on.

cs.CL↗

SlideLab: Audience-Centered Scientific Slide Generation and Evaluation

Scientific presentations are more than summaries of research papers. They need to present the work in a coherent sequence, explain the main ideas clearly, and help the audience follow the presentation. We present SlideLab, a training-free multi-agent framework for generating scientific presentations from research papers. SlideLab first plans the presentation narrative, then builds and iteratively refines a shared slide deck using agents for content planning, visual generation, layout refinement, and grounding verification. In a blind human preference study, SlideLab was preferred over both open-source and commercial systems on 77% of papers while using roughly 4 times fewer inference tokens than the strongest open-source baseline. We also introduce ConfArena, an audience-oriented evaluation framework that simulates a conference room and assesses presentations slide by slide. ConfArena matches human system rankings and detects injected presentation problems, including falsified numbers, degraded figures, dropped slides, and shuffled slide order.

cs.CL↗

Volume stability of plank covers

We prove a longstanding conjecture of András Bezdek on the stability of plank covers of the planar disk $D$. If a sufficiently small disk $\varepsilon D$ is removed, then every finite family of planks covering the resulting annulus can be rearranged to cover the whole disk. For the proof, we establish a stronger dimension-independent stability theorem. If $K \subset {\mathbb R}^n$ is an origin-symmetric convex body for $n \geq 2$, and the uncovered region lies in $\varepsilon K$ and has volume $v$, then the excess overlap is at least $c v/\varepsilon$, where $c>0$ is an absolute constant. The proof introduces a pruning procedure and a flattening mechanism. Pruning peels away curvature through local removals whose cumulative effect is guaranteed to be substantial. Flattening trades straightness for bounded multiplicity and geometrically exposes the overlap that pays for each pruning step.

math.MG↗

A Log-Star Comparison Between Expectation Threshold and Fractional Expectation Threshold

We proved a log-star comparison between the expectation threshold $q(\mathcal F)$ and the fractional expectation threshold $q_f(\mathcal F)$ for any nontrivial increasing family $\mathcal F$ on a finite ground set $V$ of size $|V|=N$. Specifically, we show that $$q_f(\mathcal F)\le64\log_2^*(N+2)\,q(\mathcal F),$$ where $\log_2^* x$ is the iterated logarithm of base-2. Using similiar methods, one can extend this result to the following stronger form: there exists a universal constant $C>0$ such that \[ q_f(\mathcal F)\le Cq(\mathcal F)\log_2^*(l(\mathcal F)),\quad q_f(\mathcal F)\le Cq(\mathcal F)\log_2^*(\operatorname{VC}(\min\mathcal F)), \] where $l(\mathcal F):=\max\{2,\max_{S\in\min\mathcal F}|S|\}$ and $\operatorname{VC}(\min\mathcal F)$ denotes the VC-dimension of $\min\mathcal F$. On October 7, 2026, OpenAI released a result proving the equivalence of fractional and integral thresholds, which supersedes the findings in this note. The authors are keeping this note for historical reference.

math.CO↗

RecToolBench: Benchmarking Recommendation-Specific Tool Orchestration under Fuzzy User Intent

Recent advances in agentic recommender systems are shifting recommender systems from passive filtering engines to instruction-following agents that use external tools to resolve user intent. However, existing benchmarks often assume explicit user intent, simplified tool environments, or isolated function calls, leaving realistic tool orchestration for recommendation underexplored. To bridge this gap, we propose RecToolBench, a Model Context Protocol (MCP)-based benchmark for evaluating tool-using recommender agents under fuzzy user instructions. RecToolBench contains more than 1,200 executable tasks across three recommendation domains, 13 MCP servers, and 32 tools, spanning single-tool calls, parallel tool calls, sequential tool chains, and hybrid tool orchestration. We construct RecToolBench with a scalable synthesize--fuzzify--judge pipeline that generates executable fuzzy recommendation tasks, and evaluates agent trajectories using rule-based execution checks and rubric-based LLM evaluation. Experiments on representative LLMs show that syntactically valid tool calls do not guarantee successful recommendations. Models struggle with semantic parameter grounding, multi-step evidence integration, and grounded final recommendations, especially as orchestration complexity increases. Our results identify tool orchestration under fuzzy user intent as a major bottleneck for agentic recommender systems. Our data and code are available at https://github.com/ShawnChenn/RecToolBench.

cs.IR↗

CCRV-Bench: Constraint-Based Evaluation of Causal Reasoning in Vision-Language Models

Vision-language models (VLMs) have demonstrated excellent performance in visual tasks, but their visual causal reasoning capabilities still lack reliable evaluation. Existing evaluations struggle to distinguish whether a model is performing causal reasoning based on visual evidence or relying on statistical correlations for shortcut learning, thereby potentially overestimating their actual capabilities. This paper proposes CCRV-Bench, a constraint-driven visual causal reasoning benchmark for single-image physical scenarios. We construct an orthogonal framework that evaluates four causal task dimensions: causal relation discovery, state prediction, causal diagnosis, and intervention-outcome prediction. We further introduce entity symbolization, spatial grounding, the factual adversarial constraint, and minimalist output constraints to reduce shortcut cues while preserving the physical commonsense required by the task. Experiments across 14 multimodal models show that constraint sensitivity is task- and model-dependent: intervention-outcome prediction has the largest average effective degradation among the four causal tasks, spatial grounding is the most damaging constraint on average, and the factual adversarial constraint improves DCR for all evaluated models. These results show that unconstrained performance does not determine constrained robustness and that a single aggregate score can obscure distinct failures in causal identification, spatial grounding, and constraint-compliant expression. CCRV-Bench provides a standardized framework for diagnosing image-grounded causal reasoning under controlled constraints. The code is available at https://github.com/0815linyuan/CCRVBench

cs.CV↗

Hilbert space in the flat limit of AdS/CFT. I. Massive particles

The flat limit of AdS$_4$ spacetime corresponds, at the level of the symmetry algebra, to the Inönü--Wigner contraction $\mathfrak{so}(2,3) \to \mathfrak{iso}(1,3)$. We describe the emergence of the Hilbert space of massive scattering states of arbitrary integer spin $s$ from unitary lowest-weight representations $\mathcal{D}(Δ,s)$ of the conformal group, using a basis of `pseudo-momentum' states first introduced by Fronsdal. These states nicely reduce to Wigner's momentum eigenstates in the limit of infinite AdS curvature radius $\ell \to \infty$, at fixed $m=Δ/\ell$. A simple conformal map relates this description to conformal fields in $\mathbb{R}^3$ familiar from radial quantization, with the interior and exterior of the unit ball respectively corresponding to outgoing and ingoing momenta. Under the contraction, and upon appropriate rescaling, the inner product on $\mathcal{D}(Δ,s)$ converges to the standard Lorentz-invariant inner product of Wigner's massive particles, so that unitarity is preserved throughout. The contraction to massless particles is briefly discussed here, and will be treated in detail in a forthcoming paper.

hep-th↗

Generalized Harmonic Measures: Synchronized Approximation, Representation and Quantitative Measure Recovery

We study synchronized boundary and kernel approximation for harmonic measure systems satisfying nested mean-value identities. Under boundary concentration along a common subsequence, joint limits yield representation on a common refinement, and gluing gives uniqueness on the minimal boundary. Neighborhood nonvanishing gives full-sequence synchronized selection. For a compact set whose removal from a bounded $C^{1,1}$ domain in dimension at least three leaves a connected domain, zero Newtonian capacity is equivalent to concentration of one normalized singular kernel, constructed from harmonic measures, on a subsequence of one relatively compact exhaustion. For such zero-capacity sets, capacitary estimates give quantitative recovery of finite signed source measures on every prescribed relatively compact exhaustion. On rectifiable curve networks in three dimensions, the logarithmic error bound is sharp for this recovery formula in the class of finite measures. Boundary splitting and finite weighted metric graphs illustrate the gluing conclusions.

math.FA↗

Evaluating Budgeted Context Projection with Unexecuted Companion Runs

Context projection can shorten individual requests while changing whether an agent finishes within its budget. We examine how sequential evaluation obscures this trade-off when a capped first continuation prevents its companion from running. In a recorded ReVerPi source-reading campaign, 15 pairs with two final answers yield 12 historically scored successes per arm. Retaining all 27 intervention boundaries distinguishes observed failures from ten unexecuted companions and bounds projected-minus-full success between $-9$ and $+1$ tasks. Under the archived scoring contract, a frozen projection selector has a success difference from full context of $[-3,0]$; outside four fitting tasks, it is $[-4,-1]$ across 23 boundaries. Excluding one task whose platform premise is not established by retained actor-input evidence changes the nonfitting range to $[-3,0]$ across 22 boundaries, so strict inferiority is not robust to that exclusion. Among eleven historically joint-success pairs, projection uses 25% fewer aggregate logical tokens but more tokens for the median pair and 55 rather than 35 suffix requests. These retrospective results concern one adaptively assembled campaign, not population performance. The case motivates accounting that retains every boundary, preserves unknown outcomes and policy dependencies, checks task premises, and separates bounded completion from success-conditioned resource use.

cs.AI↗

Parallel and Distributed Fermionic Simulation via Dynamic Encoding

We demonstrate a simple and efficient method to parallelize and distribute Trotterized Hamiltonian simulation of fermionic systems across multiple QPUs. Using combinatorial covering designs to define a minimal set of fermion-qubit encodings, we demonstrate communication cost scaling as $\mathcal{O}(M q^4 r)$ for a system of $M$ fermionic modes and Trotter number $r$, improving on the static encoding bound for $q$ QPUs, $\mathcal{O}(M^4 q r)$. We compare this approach to dynamic encoding using a randomised method, Pauli-weight based optimisation and hypergraph partitioning. Applying these to the Hamiltonians of a range of molecular systems split across two QPUs, we find the combinatorial covering approach results in the lowest communication cost in all but the sparsest Hamiltonians.

quant-ph↗

Using LMs to Model the Effects of Context and Coreference during Sentence Comprehension

Language models (LMs) are often used as a tool to model human language processing. Recent studies suggest that severely restricting LMs' context window improves their fit to human psycholinguistic data by simulating human working memory constraints. However, it is possible that this strict memory-decay approach overlooks humans' reliance on long-range structural representations, such as discourse structre. In this work, we systematically vary the context window size of GPT-2 across four large-scale naturalistic English reading-time datasets and observe a U-shaped relationship: Although restricted contexts (< 20 tokens) successfully capture local memory limitations, expanded contexts (500--1,000 tokens) ultimately yield the highest overall psycholinguistic fit. To investigate the mechanism driving this benefit, we conduct a counterfactual inference-time experiment that disrupts cross-sentential entity chains by pronominalizing repeated discourse entities. Obscuring these structural linkages significantly degrades the predictive power of larger context windows by 20% to 40%. Our experiments demonstrate that tracking long-range coreference relations is one important factor for the alignment between LM surprisal and human reading behavior, and approximate the extent to which human comprehenders use global discourse relations during language processing.

cs.CL↗

READ-Bench: Benchmarking Historical Instance Retrieval for Time-Series Diagnosis

Time-series diagnostic systems rarely rely on retrieving relevant historical cases, and when they do, retrieval is evaluated only indirectly through downstream prediction. We introduce READ-Bench, a benchmark for historical-case retrieval across 12 diagnostic datasets, centered on multivariate time series, that defines relevance by shared fault or event type rather than signal shape, so visually different traces of the same fault count as relevant while similar-looking traces of different faults do not. Treating retrieval as a base retriever followed by a reranker, we evaluate classical distances, symbolic retrievers, self-supervised and foundation-model embedders, and their fusion, plus label-aware and language-model rerankers, under one protocol that varies supervision, pollution, and corpus scale with significance testing. Under a common channel-independent interface, pretrained representations offer no statistically detectable advantage over strong classical and symbolic baselines for search alone. The decisive factor is a small amount of resolved-case supervision at reranking, namely a Gaussian-process reranker that propagates a few neighbor labels in embedding space, which helps far more than more sophisticated representations or language-model reasoning and holds under pollution and at full corpus scale. Guided by these findings, we fuse a normal-residual-scored embedder with a dynamic time warping leg via reciprocal-rank fusion, then rerank with the Gaussian-process reranker, improving NDCG@10 over its own search stage on all 12 datasets, by +0.11 from reranking and +0.16 over the strongest single base retriever.

cs.AI↗

PlurVA-LLM-2026 Shared Task Track-1: Pluralistic Value Alignment in LLMs via Multilingual Fine-Tuning and Threshold Calibration

We present our system for the PlurVA-LLM 2026 Shared Task Track-1, which focuses on pluralistic value alignment in the contexts of China, Indonesia, and Sri Lanka. For this resource-constrained track, we fine-tuned Llama 3.1 8B Instruct using 4-bit QLoRA. Our approach combines option-permutation augmentation for Chinese data, annotator vote expansion for Indonesian data, and binary reformulation with SinhalaMMLU augmentation for Sri Lankan data. We further applied conditional threshold calibration to the predictions for the Sri Lankan data. The final system achieved accuracies of 0.785 for Chinese, 0.715 for Indonesian, and 0.916 for Sri Lankan, resulting in an overall macro-average accuracy of 0.805.

cs.CL↗

FA-Bench: A Benchmark for Phone- and Word-Level Timestamp Accuracy in Forced Alignment and ASR on Clean and Noisy Speech

Forced alignment estimates the timestamps of each word, phone or character in speech given its transcript. Published comparisons normalize transcripts, split the data and match boundaries differently, so their numbers cannot be read together. We present FA-Bench, an open framework that fixes those choices once and releases the code, splits, phone mapping, text normalization and scoring script, with results published periodically. Track 1 gives every aligner the reference transcript and Track 2 gives it a recognizer's output, on the same audio, clean and degraded four ways, with 21 open models and 9 commercial APIs under a unified protocol. We score every boundary of an utterance and check the two labels beside it, so a word the recognizer missed or invented is charged. Using a tolerance-based F1 as our primary metric eliminates the 9% to 14% score inflation that standard MAE causes on recognition-dependent systems in conversational speech. We then group boundaries by position and how many adjacent words were recognized correctly, which shows where a system lost the score. We discovered systematic bias in how current systems time words, with Whisper about 150 ms early and several commercial ASR APIs over 50 ms late. Code and results are at https://github.com/olewave/fa-bench

cs.CL↗

Hypercyclic Bergman-Toeplitz operators with some harmonic symbols

Little is known about the dynamics of Toeplitz operators on the Bergman space, and even in the case of harmonic symbols, few results are available. In this paper, we study the dynamical properties of Toeplitz operators with symbols of the form $a \bar{z} + p (z)$ on the Bergman space, where $a\neq 0$ and $p$ is analytic on the closed unit disk. We obtain some necessary conditions and some sufficient conditions for such operators to be hypercyclic, weakly mixing, mixing, chaotic, or frequently hypercyclic. As an application, we obtain a complete characterization of hypercyclicity for Toeplitz operators with harmonic linear polynomial symbols. We also show that Toeplitz operators with the same symbol may exhibit different hypercyclic behavior on the Hardy and Bergman spaces.

math.FA↗

Satellites and invariants of links

An invariant $v$ of $m$-component links is called "cableable" if there exists a $k$ such that whenever a link $L'$ is obtained from a link $L=(K_1,\dots,K_m)$ by replacing each knot $K_i$ with its $(p_i,q_i)$-cable for some $p_i$ and $q_i$, we have $v(L')=(p_1\cdots p_m)^kv(L)$. The following problem is implicit in a number of papers by P. M. Akhmetiev and originates from the Arnold-Moffatt program for finding topological lower bounds for the energy of a magnetic field: Does there exist a cableable finite type invariant of links in $S^3$ which is not a function of the pairwise linking numbers? A potential solution of this problem was proposed by Akhmetiev himself, with the desired invariant defined as an analytic expression involving a magnetic field modeled on the given link, but we note that basic properties that he claimed of his invariant cannot be all true. In any case, we offer a different solution, with the desired invariant being a function of the coefficients of the Conway polynomial of the link and its sublinks. Moreover, we show that the cables can be replaced by arbitrary satellites. Much of the proof is a study of low degree coefficients of the Conway potential function $Ω_L(x_1,\dots,x_n)$ expanded as a formal power series in Conway's variables $z_i=x_i-x_i^{-1}$. We also discuss type $n$ invariants which are "cableable up to an invariant of type $n-1$", some cableable invariants which are not of finite type (particularly a certain modification of Milnor's $\barμ$-invariants), and applications to links of solenoids.

math.GT↗

When to Evict, Not What to Keep: Draft-Guided Eviction for Training-Free KV-Cache Compression

Training-free KV-cache compression methods such as SnapKV, H2O, and PyramidKV evict tokens at the end of prefill, aiming to preserve the attention mass that future queries are expected to use -optimizing what to keep. We show that this objective fails in two distinct ways. (1) Compensation: restoring the evicted attention mass can recover the attention-level target without recovering task quality. (2) Selection: covering more of the true decode-query mass can hurt quality when the recovered mass is fragmented rather than concentrated in coherent spans. These failures share a common cause: eviction occurs before the queries that determine the answer trajectory exist. We propose Draft-Guided Eviction (DGE), which defers eviction until after drafting the first k=2 answer tokens using the full cache - just one decode step beyond prefill. Because the draft is generated from the answer's own prefix, no cache entries are discarded before this trajectory signal becomes available. The per-head cache budget remains unchanged, and DGE can be applied directly to SnapKV, PyramidKV, H2O, and StreamingLLM without modifying their eviction scores. Unlike extra-pass methods, DGE changes when eviction occurs rather than what cache entries are selected. Extensive experiments demonstrate that DGE outperforms prior methods at every evaluated budget on five of six instruct-tuned backbones, achieving 44.2 on LongBench, nearly matching FullKV at 44.3. The timing-only control DGE-W achieves the same score, demonstrating that the gain comes from when eviction occurs rather than what is selected - an effect we term trajectory anchoring.

cs.LG↗