Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Retrieval-Augmented Generation Must Move Beyond Factual Grounding to Represent Diverse Opinions

Retrieval-Augmented Generation (RAG) systems are built on an unexamined assumption - that queries have correct answers and retrieval should converge toward them. This position paper argues that this creates a factual bias where RAG systems optimize for reducing epistemic uncertainty while ignoring the aleatoric uncertainty, inherent in opinion-rich content. The consequences go beyond technical limitations- due to risk of minority voice erasure and risk of opinion manipulation. To address this, we formalize opinion-aware retrieval through uncertainty quantification and derive a unified objective using the Wasserstein distance. As an existence proof, we present Opinion-Aware RAG (O-RAG), which enriches documents with LLM-extracted, entity-linked opinion metadata before indexing. Across e-commerce seller forums and public hotel reviews, O-RAG reduces Wasserstein distance to corpus-level sentiment distributions by 18-48%, and human evaluators preferred its responses 79.2% of the time. We close with a research agenda for opinion-aware RAG.

cs.AI↗

Daycare Matching with Siblings: Social Implementation and Welfare Evaluation

In centralized matching markets, agents may value joint assignment, as with siblings or couples. Standard preference estimation ignores such complementarities, complicating welfare analysis of priority rules for paired assignment. We develop an empirical framework incorporating these preferences and apply it to Japanese daycare assignment. Families face both additional commuting distance and a fixed disutility from split assignment. We estimate the latter at 4.61 commuting-kilometer equivalents. Our fixed-report counterfactual estimates that the reform increased mean welfare by 0.032 kilometer-equivalent units. Ignoring sibling complementarity understates welfare gains for households applying simultaneously for multiple children by about 27%.

econ.GN↗

Gradient Descent's Last Iterate is Often (slightly) Suboptimal

We consider the well-studied setting of minimizing a convex Lipschitz function using either gradient descent (GD) or its stochastic variant (SGD), and examine the last iterate convergence. By now, it is known that standard stepsize choices lead to a last iterate convergence rate of $\log T/\sqrt{T}$ after $T$ steps. A breakthrough result of Jain et al. [2019] recovered the optimal $1/\sqrt{T}$ rate by constructing a non-standard stepsize sequence. However, this sequence requires choosing $T$ in advance, as opposed to common stepsize schedules which apply for any time horizon. Moreover, Jain et al. conjectured that without prior knowledge of $T$, no stepsize sequence can ensure the optimal error for SGD's last iterate, a claim which so far remained unproven. We prove this conjecture, and in fact show that even in the noiseless case of GD, it is impossible to avoid an excess poly-log factor in $T$ when considering an anytime last iterate guarantee.

math.OC↗

Timescale Separation Enables Deep Reinforcement Learning Control of Rotating Detonation Engine Mode Transitions

Rotating detonation engines (RDEs) are a promising propulsion concept that may offer higher thermodynamic efficiency and specific impulse than conventional systems, but nonlinear phenomena, including transitions to oscillatory or chaotic propagation modes, can hinder practical operation. Deep Reinforcement Learning (DRL) has emerged as a promising method for controlling complex nonlinear dynamics such as those observed in RDEs. However, the multi-timescale nature of the RDE system makes direct application of DRL challenging. We address this challenge by reformulating the DRL problem in a moving reference frame that follows the detonation-wave pattern, making the wave structure appear quasi-steady to the agent. This reformulation enables scale separation between fast detonation propagation and slower operating-mode dynamics. We train DRL controllers to modulate spatially segmented injection pressure in a one-dimensional reduced-order RDE model and induce rapid transitions between different mode-locked states. Across a range of actuation periods, initial states, and target modes, controllers trained in the moving frame learn more reliably than those trained in a stationary frame and remain effective over a broader range of actuation periods. These results suggest that symmetry-aware moving reference frame formulations may be useful for related multiscale flow-control problems and that scale separation should be exploited whenever possible to enable DRL control of multi-timescale systems.

physics.flu-dyn↗

The $Λ_cΣ_c$ dibaryon and $\barΛ_cΣ_c\pm \barΣ_cΛ_c$ baryonium states via the QCD sum rules

In the present work, we study the $Λ_cΣ_c$ dibaryon and $\barΛ_cΣ_c\pm \barΣ_cΛ_c$ baryonium states via the QCD sum rules. We construct four (eight) currents with definite $J^P$ ($J^{PC}$) to interpolate the dibaryon (baryonium) states and obtain twelve QCD sum rules. For the dibaryon states, the state with the $J^P=1^+$ lies below the $Λ_cΣ_c$ threshold, and is a molecule candidate. For the baryonium states, the states with the $J^P=0^{-\pm}$ and $1^{-\pm}$ lie below the $\barΛ_cΣ_c$ threshold, and are four molecule candidates. The present predictions can be confronted to the experimental data in the future and make contributions to the exotic spectroscopy.

hep-ph↗

Thermal conductivity tuning of scalable nanopatterned silicon membranes measured with a three-probe method

Phononic silicon structures have emerged as an integrable and scalable nanosystem for tailoring thermal transport. However, their widespread adoption has been limited by their complex fabrication pathways. Alongside, the reliable characterization of thermal properties in suspended nanostructured films remains challenging, as thermal contact resistances often hinder the accuracy of measurements. In this work, we demonstrate a clear and controllable reduction of thermal conductivity in nanopatterned silicon membranes. A block copolymer self-assembly approach is employed to fabricate nanoholed silicon films with a pitch of 63 nm and hole diameters of 35 nm. Additionally, we introduce an extension of the three-probe technique that enables robust, quantitative, and spatially resolved thermal conductivity measurements in complex thin-film systems, accounting for thermal contact artifacts. The method is validated through measurements on unpatterned 40 nm-thick silicon thin films between 30 and 350 K, yielding a room-temperature thermal conductivity of 46.5 W/m.K. Finally, we further show that controlled etching of the nanoholes provides a powerful means to tune thermal transport in the overall studied temperature range, establishing hole etch depth control as an effective parameter in phononic silicon. Specifically, a fivefold reduction in thermal conductivity is achieved, reaching 7.3 W/m.K for fully etched-through membranes at room temperature.

cond-mat.mes-hall↗

IQP circuits for 2-Forrelation

The $2$-Forrelation problem provides an optimal separation between classical and quantum query complexity and is also the problem used for separating $\mathsf{BQP}$ and $\mathsf{PH}$ relative to an oracle. A natural question is therefore to ask what are the minimal quantum resources needed to solve this problem. We show that $2$-Forrelation can be solved using Instantaneous Quantum Polynomial-time ($\mathsf{IQP}$) circuits, a restricted model of quantum computation in which all gates commute. Concretely, signed $2$-Forrelation can be solved by a classical random choice between two one-query $\mathsf{IQP}$ circuits, while the absolute-value variant uses two independent executions of this randomized procedure. This answers a recent open question of Girish (STOC 2026) on the power of commuting quantum computations. For the Raz-Tal distribution, this randomization is unnecessary. We use this to show that there is an oracle $O$ such that $\mathsf{IQP}^O \not\subseteq \mathsf{PH}^O$, strengthening the result of Raz and Tal (STOC 2019). It also yields an oracle separation between $\mathsf{IQP}$ and $\mathsf{DQC}_1$. We prove Fourier growth bounds for multi-query $\mathsf{IQP}$ circuits, including bounds in terms of the size of their accepting set. Our results suggest a possible route toward decision-based quantum advantage within the restricted $\mathsf{IQP}$ model. The key ingredient is an algebraic identity of the quadratic function $Q(x) = \sum_{i < j} x_ix_j$ that allows extracting inner-product phases within an $\mathsf{IQP}$ circuit.

quant-ph↗

Evaluating a Layered Prompt-Injection Defence for the Model Context Protocol: A Record-Level Audit of Decision Conventions, Corpus Provenance and Reproducibility

The Model Context Protocol (MCP) extends the prompt-injection attack surface of large language model applications to tool descriptions, parameter schemas and tool outputs. Defences for it report detection figures that are not comparable, because each is measured on its authors' own corpus under a decision convention that is rarely stated. This paper audits one such evaluation at the level of individual decision records. Its object is CASCADE, a fully local layered defence (rule matching, embedding similarity and an optional local language-model review), run in three configurations on a frozen 5,000-sample corpus under a pinned code revision; the corpus, the per-sample records and the analysis scripts are public. Four findings result. First, the convention that collapses allow, review and block onto a binary label sets the headline: the pipeline without review reports an 11.70% false-positive rate when referrals count as positives and 1.51% when only denials do, while referring 68.5% of all traffic to a human. Second, the construction of the corpus shapes the aggregate: 65.6% of records are texts placed in one of five fixed wrappers, the wrapper alone identifies the label, and for the same text wrapping lowers the false-positive rate from 21.2% to 3.2% and raises detection from 86.0% to 98.0%. Third, the published description does not identify what ran: the rule module inside the pipeline is not the rule engine evaluated alone; the block threshold produced denials only among the first 48 requests; and, without review, every later denial, including all 23 denials of benign requests, came from an output guard run on an echo of the input. Fourth, the local review model, invoked on 32.6% of requests at 2.5 s each, changes no binary outcome but converts 1,492 referrals into denials. A ten-item reporting checklist is derived from these findings.

cs.CR↗

On a relation of a conjecture of Goncharov to the co-Lie algebra of Bloch-Kriz mixed Tate motives

Goncharov defined for each field $F$ and an integer $n$ greater than 1 a certain group $B_n(F)$. We consider the possibility of defining a linear map from $B_n(F)$ to the co-Lie algebra of the category of mixed Tate motives defined by Bloch and Kriz, in terms of motivic polylogarithms. We give results which support this possibility assuming part of the conjecture by Beilinson and Soulé on vanishing of $K$-groups of fields.

math.AG↗

Three-dimensional time-periodic problem on the Boltzmann equation with external force

The time-periodic problem for the Boltzmann equation with a prescribed time-periodic external force in the three-dimensional whole space has remained open since it was first treated in [15], where the result was restricted to spatial dimensions not less than five. This paper gives an affirmative answer when the force is sufficiently small in $\mathcal{C}(\mathbb{R};\dot{B}^{-3/2}_{2,\infty}\cap\dot{H}^N)$ with $N\geq 4$. The proof follows Serrin's method, through the global stability of the Cauchy problem for a time-periodic force. As a direct consequence, the same argument yields the existence and stability of stationary solutions in three dimensions for a time-independent force, which is not required to be a gradient and may therefore be rotational.

math.AP↗

How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks

The wide adoption of AI agents in complex human workflows is driving rapid growth in LLM token consumption. When agents are deployed on tasks that require a significant amount of tokens, three questions naturally arise: (1) Where do AI agents spend the tokens? (2) Which models are more token-efficient? and (3) Can agents predict their token usage before task execution? In this paper, we present the first systematic study of token consumption patterns in agentic coding tasks. We analyze trajectories from eight frontier LLMs on SWE-bench Verified and evaluate models' ability to predict their own token costs before task execution. We find that: (1) agentic tasks are uniquely expensive, consuming 1000x more tokens than code reasoning and code chat, with input tokens rather than output tokens driving the overall cost; (2) token usage is highly variable and inherently stochastic: runs on the same task can differ by up to 30x in total tokens, and higher token usage does not translate into higher accuracy; instead, accuracy often peaks at intermediate cost and saturates at higher costs; (3) models vary substantially in token efficiency: on the same tasks, Kimi-K2 and Claude-Sonnet-4.5, on average, consume over 1.5 million more tokens than GPT-5; (4) task difficulty rated by human experts only weakly aligns with actual token costs, revealing a fundamental gap between human-perceived complexity and the computational effort agents actually expend; and (5) frontier models fail to accurately predict their own token usage (with weak-to-moderate correlations, up to 0.39) and systematically underestimate real token costs. Our study offers new insights into the economics of AI agents and can inspire future research in this direction.

cs.CL↗

On proper compactifications of topological groups

In the present paper, we examine in detail the method of "graph compactifications" of topological groups. The graph and Ellis methods of constructing proper compactifications of topological groups are applied for the investigation of possible extensions of algebraic operations on a topological group to its compactifications, and give descriptions of Roelcke, Ellis, WAP, and graph compactifications of topological groups. Additionally, using dichotomy theorems of A.V.Arhangelskii, we show that the description of compactifications can be effectively used in the investigation of topological properties of their remainders. As examples, subgroups of the permutation group (in the permutation topology) and the automorphism group of a LOTS (in the topology of pointwise convergence) are examined.

math.GN↗

Move-Then-Operate: Behavioral Phasing for Human-Like Robotic Manipulation

We present Move-Then-Operate, a Vision language action framework that explicitly decouples robotic manipulation into two distinct behavioral phases: coarse relocation (move) and contact-critical interaction (operate). Unlike monolithic policies that conflate these heterogeneous regimes, our architecture employs a dual-expert policy routed by a learnable phase selector, introducing a structural inductive bias that isolates phase-specific dynamics. Phase labels are automatically generated via an MLLM-based pipeline conditioned on lightweight contextual cues such as end-effector velocity and subtask decomposition to ensure alignment with human motor patterns. Evaluated on the RoboTwin2 benchmark, our method achieves an average success rate of $68.9\%$, outperforming the monolithic $π_0$ baseline by $24\%$. It matches or exceeds models trained on $10\times$ more data and reaches peak performance in $40\%$ fewer training steps, demonstrating that architectural disentanglement of move and operate phases is a highly effective and efficient strategy for mastering high-precision manipulation.

cs.RO↗

Point-MF: Stabilizing One-Step Mean Flows for Single-Image Point Cloud Reconstruction

Single-image point cloud reconstruction requires recovering complete object-level geometry, including occluded regions, from a single RGB image. Diffusion- and flow-based reconstructors can model this ambiguity, but their iterative sampling requires many network function evaluations, while naive one-step point-space updates often produce outliers and density imbalance. We propose Point-MF, a one-step conditional Mean-Flow framework for geometry-only point cloud reconstruction. Point-MF predicts an interval-averaged velocity field directly in point-cloud space and reconstructs the output with a single network function evaluation, without training a point-cloud autoencoder. To stabilize large Mean-Flow jumps, we introduce Denoised Space Anchor (DSA), a set-distance auxiliary loss that anchors the denoised point set implied by the predicted velocity to the ground-truth geometry. On all 13 ShapeNet-R2N2 categories and Pix3D, Point-MF achieves strong reconstruction quality under the aligned point-set protocol, improving average CD and EMD over the evaluated point-cloud reconstruction baselines. It runs in 40.72 ms per sample on an RTX A4000, within the same latency order as the feedforward RGB2point baseline and substantially faster than iterative diffusion baselines such as PC$^2$ and BDM. These results show that denoised-space geometric anchoring enables stable one-step Mean Flow directly in point-cloud space.

cs.CV↗

Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all---they are failures of the benchmark itself: broken specifications, implicit assumptions, and rigid evaluation scripts that penalize valid alternative approaches. We propose employing frontier LLMs as systematic auditors of evaluation infrastructure, and realize this vision through BenchGuard, the first framework explicitly designed for joint cross-artifact auditing of execution-based agent benchmarks. BenchGuard cross-verifies all benchmark artifacts via structured LLM protocols, optionally incorporating agent solutions or execution traces as additional diagnostic evidence. Deployed on two prominent scientific benchmarks, BenchGuard identified 12 author-confirmed issues in ScienceAgentBench---including fatal errors rendering tasks unsolvable---and exactly matched 83.3% of expert-identified issues on the BIXBench Verified-50 subset, catching defects that prior human review missed entirely. A full audit of 50 complex bioinformatics tasks costs under USD 15, making automated benchmark auditing a practical and valuable complement to human review. A preliminary native-format audit of ProgramBench further demonstrates cross-format applicability. These findings point toward AI-assisted benchmark development, where frontier models serve not only as subjects of evaluation but as active participants in validating the evaluation infrastructure itself.

cs.CL↗

Galilean Reeh-Schlieder obstruction in the vacuum and the thermal Reeh-Schlieder property

We ask in which states, and for which local algebras, Galilean quantum fields have the Reeh-Schlieder property. The nets are generated by time-zero canonical fields on spatial regions and by the dynamics. The Bargmann mass over the charge of the field is a number operator for the time-zero fields, bounded from below by positivity of the energy and boost covariance, and Chaiken's theorem gives Galilean Fock rigidity: these fields form a direct sum of Fock representations, and the time-zero annihilation fields annihilate every translation-invariant vector. If the vacuum is the only vector invariant under spatial translations (not implied by the usual uniqueness axiom), the vacuum sector of the observable net is one-dimensional. No vector of finite mass is cyclic or separating for an algebra of a spatial region at fixed time. In the vacuum of Bose gases with bounded or Coulomb-type pair potentials (with Buchholz's renormalised dynamics if not H-stable), the fields of any region over any time interval generate all bounded operators. By contrast, a KMS state is cyclic and separating for the algebra of every region $B\times I$ with temporal extent whenever the spatial algebras at all times generate the global algebra, as for free gases. For the free Schrödinger field in the vacuum, the thermal states of Bose and Fermi gases and the Fermi sea, fixed-time field algebras are type I factors, for which thermal vectors and the Fermi sea are separating but not cyclic, and every field algebra with temporal extent is the global one: type I in pure states, type III$_1$ in thermal states, with the grand-canonical dynamics as modular group. Galilean vacua are thus never separating for local algebras. For the free field the type of the algebras of regions with temporal extent depends on the state, and what is specifically relativistic is the locality of type III$_1$.

quant-ph↗

Predictive Spatio-Temporal Scene Graphs for Semi-Static Scenes

We have seen tremendous recent progress in our ability to build "spatio-semantic" representations that enable robots to perform complex reasoning across geometry and semantics. However, the vast majority of these methods lack any ability to perform reasoning across time. This is a desirable property in situations where a robot repeatedly observes an environment where instances may change in between observations, but in a structured way. Consider as an example a home environment where the location of a mug typically moves from the cupboard to a countertop to the sink and then back to the cupboard on a daily basis. We should be able to learn this cyclic behavior and use it to predict the state of the mug in the future. In this work, we propose a method that is able to perform this type of tempo-spatio-semantic reasoning. Underpinning the method is a filter, Perpetua*, that performs Bayesian reasoning on the states of the environment that are observed over time. This filter is integrated within a 3D scene graph structure that we call PredictiveGraphs, where nodes represent objects and edges function as Perpetua* filters encoding spatio-semantic relationships. We validate the method in both simulation and real-world dynamic navigation tasks, where our real-world experiments consist of an environment that is undergoing semi-static changes at a bi-hourly frequency over a period of three weeks. In both settings, we demonstrate that our method outperforms baselines in predicting future environment states, even in the presence of distributional shifts.

cs.RO↗

Useful Features, Backward Scores: OOD in Language-Model Trajectories

Out-of-distribution (OOD) detectors prioritize inputs for closer inspection. Yet features that distinguish input groups need not yield a useful anomaly ranking. We analyze this gap in language-model trajectories under text-length control and fixed score directions. On Spam development data, an input adaptation of D^2HScore falls from raw AUROC 0.919 to 0.530 after length matching. On length-matched, held-out HateSpeech inputs, the same features yield AUROC 0.644 for a labeled linear classifier but 0.444 for an ID-fitted distance score. ToxicChat shows the same contrast. Feature-selection and backbone controls retain the main reversal pattern. Frozen Civil Comments and TweetEval irony tests also reverse (0.467 and 0.435), extending the finding beyond toxicity. In these contrasts, anomalous groups have farther centers but tighter spread. A labeled, fixed-center feature-space intervention changes rankings: equalizing spread helps some tasks and harms others. OOD evaluation must check the chosen score's ranking even when its features distinguish the classes.

cs.CL↗