Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 703 records · Page 39Linked to original sources

Squeezed light from a semiconductor amplifier

The possibility of generating light in a squeezed quantum state during saturated light amplification in a semiconductor amplifier is theoretically analyzed. Due to four-wave mixing process during amplification, weak components of the spontaneous emission of the medium with symmetric frequency shifts relative to the strong amplified wave acquire the quantum properties of two-mode squeezed light. The squeezing of quadrature components occurs in the range of frequency shifts of several GHz ($\sim$ the inverse relaxation time of the charge carrier density) and can reach 10 dB or more under optimal conditions. The formation of a two-mode squeezing phase facilitates the observation of the squeezing effect, in which a strong wave, playing the role of a local wave, exhibits quadrature noise suppression without additional homodyne detection.

quant-ph↗

Parameterized Reachability for Register Machines with Data

We investigate the parameterized reachability problem for concurrent register machines over infinite data domains. In this framework, each machine is a program equipped with a set of local registers, the communication across machines is mediated through a set of shared registers. Both local and shared registers can take values from an infinite data domain. The program's primitive operations include copying values between registers, assigning constants, comparing registers for (dis-)equality, and nondeterministic assignments that store an arbitrary domain value into a local register. The parameterized reachability problem considers a program and a target location, asking whether there exists some n in Naturals such that an execution of n identical machines (referred to as instances) results in at least one instance reaching the specified location. We show that this problem is Pspace-complete in the general case and it becomes undecidable if a freshness assumption (i.e., each assignment must produce a unique value distinct from all constants) is applied to nondeterministic assignments. This undecidability persists even for systems restricted to two shared and two local registers. Finally, we establish optimal decidability results for two restricted settings: when each thread is limited to a single local register, or when the system utilizes only one shared register.

cs.FL↗

Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight

Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails. We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting? We introduce prospective learning, which uses post-hoc experience to supervise foresight predictions from the pre-interaction view, and instantiate it with Self-Retrospection Distillation (SRD). Intuitively, a completed trajectory reveals knowledge that would have been useful and pitfalls that should be avoided; SRD distills this privileged hindsight into trajectory-blind foresight of the same policy. Foresight serves only as a training target and need not be explicitly generated at inference time. Across 10 tool-integrated reasoning and long-horizon agentic tasks, SRD complements RLVR and self-distillation baselines with gains of up to 24.2 pp. Its advantage is especially pronounced when reward contrast is scarce: when 37--98% of rollout groups are reward-uniform across model scales, yet SRD can still exploit learning signal from sampled trajectories. In the 2B setting, where 98% of groups are all-failure, the RLVR training ends up at 0.0% success, while adding SRD reaches 60.6% under the same rollout budget. Our results suggest that post-hoc agent experience is useful not only for evaluating or improving behavior, but also for shaping predictive representations before available interaction.

cs.AI↗

When Plans Change Answers: Formalizing Cost-Accuracy Optimization for Semantic Queries

In semantic query engines, predicates are evaluated by machine-learned models, and the choice of a query plan affects not only the cost of a query but also its result. Existing systems either apply a fixed threshold to each semantic operator or tune accuracy per operator, without accounting for how errors propagate through joins. We give a formal problem definition for cost-accuracy optimization of such queries. Our starting point is the calibrated confidence that decision models such as Jev attach to each decision. It yields an expected error for every decision; weighting these errors by each decision's contribution to the output (in the simplest case, its fan-out) gives the expected output quality of a plan without any labeled data, and the same computation in reverse turns an output-level accuracy target into a price on each base or intermediate tuple. Building on this, we define an oracle semantics for relational algebra with semantic operators, physical plans as pairs of a logical plan and a decision policy, declarative output-level targets, and a hierarchy of plan equivalence. We show that accuracy is plan-invariant under pointwise-deterministic policies, and that selection pushdown is not quality-sound when escalation bands are calibrated on the plan's own candidates. Expected quality can be computed in polynomial time under bag semantics; under set semantics it follows the dichotomy of tuple-independent probabilistic databases when every relation carries a semantic predicate. Choosing which tuples to drop is NP-hard, while the optimization problem decomposes into per-tuple decisions through two Lagrange multipliers, and, with what we call confidence-centric skipping, tuples that can no longer affect the target are skipped without being scored. Simulations on a synthetic workload illustrate these effects; an evaluation on real engines is left for future work.

cs.DB↗

VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models

Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but incur high computational costs from processing long token sequences at every control step, limiting real-time deployment. Visual token pruning offers a direct solution, as visual patches dominate the input sequence and contain considerable redundancy. Existing approaches, however, either rely on indirect training-free heuristics, such as attention scores and motion thresholds, or require costly fine-tuning of the base VLA model. We introduce VLA-ACL (Action Consistency Learning), which learns a lightweight visual token pruning policy through action-level supervision while keeping the base VLA model entirely frozen. The training objective encourages actions produced from pruned visual contexts to remain consistent with the full-context teacher, with ground-truth actions as auxiliary supervision. This directly ties token selection to its effect on the downstream control output. Experiments on LIBERO and real-world manipulation tasks show that VLA-ACL prunes up to 87.5% of visual tokens while retaining competitive performance, reduces computation by up to 75%, and achieves a 1.5x inference speedup. These results establish a stronger performance-efficiency trade-off than existing frozen-VLA pruning methods and demonstrate the value of action-level supervision for visual token selection. Code is available at https://github.com/du-owen/VLA-ACL.

cs.RO↗

A duality-preserving extension of the Worley-Sagan insertion and Haiman's mixed insertion for the hyperoctahedral group

The Worley-Sagan insertion and Haiman's mixed insertion are insertion algorithms for shifted Young tableaux, and each of them gives a Robinson-Schensted-type correspondence between the symmetric group of degree $n$ and a set consisting of certain pairs of same-shape shifted Young tableaux with $n$ cells. It is a known fact that these two insertions are dual to each other. Our purpose is to give an extension of these two insertions without losing the duality relationship. The extended ones will be insertions producing pairs of shifted tableaux from colored permutations. Our extension of the Worley-Sagan insertion is different from the restriction of Sagan's own "Knuth version" to colored permutations. In proving the duality between our extended insertions, we "embed" them into Shimozono and White's doubly mixed insertion for unshifted tableaux by "doubling" shifted tableaux and use the self-duality of the doubly mixed insertion shown by Shimozono and White.

math.CO↗

Testing epimorphism onto the bicyclic monoid is $\mathsf{NP}$-complete

We prove that deciding whether there is a surjective homomorphism from an arbitrary finitely presented inverse monoid onto the bicyclic monoid is $\mathsf{NP}$-complete. As part of the proof, we show that an extension of existential Presburger arithmetic which involves greatest common divisors on $n$ arguments is in $\mathsf{NP}$, extending a recent result of Défossez, Haase, Mansutti, and Pérez (SODA 2024).

math.GR↗

Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?

Learning from experience is essential for LLM agents to adapt to unfamiliar and dynmaic environments. Evaluating this ability is therefore important for understanding how effectively agents acquire and use new knowledge. Existing benchmarks have sought to evaluate this ability, but they primarily evaluate tasks whose rules are provided in the instructions or already familiar to pretrained models, making it difficult to distinguish learning from interactions from reasoning with existing knowledge. To address this, we introduce Learn2Play Bench, a benchmark of newly designed text-based games, whose rules are novel or counterintuitive, requiring agents to acquire knowledge through interaction rather than rely solely on pretrained knowledge. These games provide reproducible feedback and automatic scoring, enabling controlled evaluation of learning across repeated attempts. We also vary game instances to test whether agents can apply what they have learned to new situations. Therefore, we evaluate how backbone models, self-evolving methods, and agent harnesses affect agents' learning ability, revealing three findings: (1) Experience retention: Retaining complete records of actions and feedback can support more effective learning than summarizing these experiences into rules or strategies. (2) Human agent gap: Top-performing human players achieve higher peak scores than the evaluated agents. Human explore more varied strategies, and repeat actions less. (3) Harness matters: With the backbone fixed, changing the harness can improve performance while reducing estimated inference cost. Together, these findings provide insights into how LLM agents learn from experience and suggest directions for future work to improve their learning ability. Project website: https://liushiliushi.github.io/learn2play-bench-website/

cs.AI↗

Thermal Transmission in a Film-Bulk Domain: Finite-Thickness Film and Thin-Film Limit

We investigate a thermal transmission problem in a domain composed of a film region and a bulk region coupled through an interface where temperature continuity and heat-flux balance are imposed. The analysis addresses both the finite-thickness film configuration and the corresponding thin-film limit, with anisotropic, non-homogeneous thermal conductivity tensors. Letting the film thickness vanish directly leads to a degenerate problem due to the vanishing measure of the film region. To preserve the contribution of the film in the limit, we introduce suitable scaling assumptions for the normal heat fluxes, namely for the flux balance across the film-bulk interface and for the flux prescribed on the external boundary of the film region, together with the requirement that the volumetric source in the film be uniformly bounded with respect to the film thickness. Under these scaling assumptions, we derive a non-trivial thin-film limit formulation that takes the form of a coupled bulk-interface problem, where the effect of the collapsing film is represented through an additional balance relation posed on the interface. Well-posedness results are established for both the finite-thickness film and thin-film limit configurations, and a rigorous convergence analysis is carried out to justify the asymptotic limit. The main mathematical achievement consists in finding explicit convergence rates in suitable functional norms, relying only on the weak regularity of the finite-thickness problem due to the transmission conditions at the film interface. To complement the study, analytical solutions are derived for a simplified film-bulk geometry and are used to compare the thermal responses of the two configurations. These solutions also provide a quantitative assessment of the convergence rates established in the theoretical analysis.

math.AP↗

PolarScale: A Physics-Grounded Benchmark for Radiometrically Consistent RGB-to-Stokes Estimation

Polarization imaging provides physical cues beyond intensity imaging but typically requires specialized hardware. Recent methods infer polarization from RGB-like inputs, yet predict only normalized Stokes components or relative descriptors, from which the radiometric scale needed for full Stokes reconstruction has been divided out. We introduce PolarScale, a benchmark that makes this scale an explicit prediction and evaluation target. Built on existing trichromatic full-Stokes measurements, PolarScale takes the per-scene normalized total-intensity image $s_0$ (a scene-referred linear image, not a consumer sRGB photograph) and asks models to predict normalized Stokes components, AoLP/DoLP/DoCP, and a per-scene scale. Because the scale is divided out of the input, it is not physically identifiable; PolarScale therefore evaluates dataset-conditioned semantic scale estimation against a constant-scale control, together with angular, self-consistency, and physical-bound metrics. Across seven restoration-based and generative backbones and three prediction strategies, the strongest restoration models estimate the scale with 3.6-4.3% mean relative error versus 5.7% for the constant control and violate physical bounds on fewer than 0.25% of pixels, whereas two generative baselines collapse to a near-zero scale; explicit descriptor supervision improves descriptor accuracy (23.66 vs. 18.88 dB PSNR for MAE). Predicted full-Stokes representations improve diffuse/specular separation, material segmentation, and glare classification, although in diffuse/specular separation the learned scale performs only on par with the constant control.

cs.CV↗

Quasi-quantum groups arising from cotensor coquasi-Hopf algebras

In this paper, we show that for any coquasi-Hopf algebra $H$ and any $H$-Majid bimodule $M$, the cotensor coalgebra $\operatorname{CoT}_H(M)$ admits a graded coquasi-Hopf algebra structure. We then define the standard filtration of a coquasi-Hopf algebra and show that the associated graded coquasi-Hopf algebra admits an embedding into a cotensor coquasi-Hopf algebra that preserves degrees zero and one. If, in addition, this coquasi-Hopf algebra is coquasitriangular (respectively, coribbon), then the cotensor coquasi-Hopf algebra is also coquasitriangular (respectively, coribbon), and the above embedding is a morphism of coquasitriangular (respectively, coribbon) coquasi-Hopf algebras. As a byproduct, we prove that the coquasi-antipode of a coquasi-Hopf algebra with the dual Chevalley property is bijective. Moreover, we establish graded quasi-coquasi-Hopf pairings between tensor quasi-Hopf algebras and cotensor coquasi-Hopf algebras.

math.QA↗

Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations

Frontier AI evaluations increasingly use open-ended, agentic, long-horizon tasks whose transcripts can span hundreds of pages of outputs and actions from complex multi-agent networks. The observability envelop-the range of what evaluators can reliably infer about an agent's behaviours-is therefore narrowing. Language model assistants can help classify and interpret agent behaviour but also afford human evaluators significant analytical degrees of freedom, threatening the reproducibility and auditability of language-model-based transcript analysis. Transect is an open source package built on Inspect Scout to help evaluators understand how a long agent run unfolded, identify behaviour worth investigating, and check interpretations against the transcript. Users specify task context and behavioural vocabulary in a reusable evaluation-family configuration, with judge models and analysis settings supplied separately. Transect's navigable reports align recorded events, token use, sub-agent activity, and model-generated behavioural labels on a common turn-based timeline. Reviewers can quickly grasp a run's narrative, trace any label or event to its source turns, and export the underlying data tables for cross-run analysis. We demonstrate the workflow on an AI R&D evaluation that generated almost 13 million tokens, dividing the agents' work into behavioural phases aligned with research-skill classifications, sub-agent delegations and interactions, and token use. The combined view shows a focus on operational work and manuscript production, with little evidence of a sustained hypothesis generation stage-arguably a necessary component for high-quality scientific outputs. Transect's flexible, customisable transcript-analysis pipeline will enable evaluators to keep pace with longer, more complex, more frequent AI evaluations while supporting scientific rigour, transparency, and reproducibility.

cs.AI↗

A Stevens's Power Law Check-up of GPT-5.5's Implicit Reading of Visual Encoding

We adapt Stevens's power law to measure the implicit ability of AI models to read visualizations, which can reveal the built-in perceptual mechanisms of algorithmic models. In our pilot study, AIs see no legend. In the color conditions, no colormap name is provided either. GPT-5.5 first views a reference visual representation and estimates its magnitude, then estimates the magnitude of each subsequent image of the same representation relative to that reference. Our evaluation of twelve visual variables makes how algorithmic models read visual encodings measurable, comparable with human perception, and more transparent to humans.

cs.CV↗

One-Shot Private Confidence Regions via Resampling

We propose a simple framework for constructing differentially private confidence regions \textit{in one shot}, i.e., by adding noise only to the final resampling quantile instead of privatizing the estimator computed on each resample. The cost of privacy of our procedure is only logarithmic in the number of resamples $B$ under with-replacement ($m$-out-of-$n$) sampling and independent of $B$ under without replacement sampling (subsampling), avoiding the $\sqrt{B}$ factor that arises in previous works. We provide nonasymptotic Gaussian Differential Privacy (GDP) and utility guarantees for both subsampling and $m$-out-of-$n$ resampling, covering mean-like estimators with small global sensitivity as well as estimators admitting efficiently computable smooth sensitivity bounds, including quantiles and degenerate U-statistics. This allows us to also obtain private confidence regions for degenerate U-statistics where the private error is much smaller than the non-private error. In all, we provide a toolbox for widely applicable DP uncertainty quantification procedures under popular resampling strategies while avoiding the computational and privacy costs of privatizing many intermediate resample statistics.

stat.ML↗

Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness

Streaming video-language models must decide not only what to answer, but whether the evidence needed for the current question has arrived. Existing systems learn that decision as a separate trigger; we ask whether an unmodified model already computes it. We show that frozen VideoLLMs carry a linearly readable evidence-readiness signal, labelled from timestamped evidence rather than from model output. It decodes in all seven models of a shared byte-identical evaluation (AUROC 0.733-0.905 under the strictest not-ready sampling, where a fitted clock is near chance), and a probe fitted without any of a benchmark family's footage still reads that family. It is question-conditioned: on byte-identical windows, changing only the question reverses the readout on 66.1% of pairs, while every question-blind control is at chance by construction. The model can answer incorrectly and still encode readiness: AUROC remains 0.722 among wrong answers. Readiness also beats uncertainty estimators and their supervised combination on latency-matched answer selection, and tracks independent human judgments more closely than confidence. Released streaming triggers are also linear readouts, yet a trained trigger read on its own base model's activations is approximately orthogonal to readiness and decodes it far less accurately than a probe. We turn the readout into Readiness Gating, an answer-timing policy that improves accuracy by up to +9.75 pp at matched video duration with negligible computational overhead. How much it gains varies with the accuracy headroom the task makes available: across 26 configurations the gain tracks that headroom, and an intervention that moves it over identical pixels moves the gain with it.

cs.CV↗

Connected Geometries in Gravitational Path Integrals, and Normalized No-Boundary Probabilities

We argue that some of the puzzling features associated with the normalization of gravitational wave functions can be substantially alleviated by restricting the path integral to a sum over connected, allowable geometries. This prescription gives sensible results when applied to classical transitions and it provides non-trivial no-boundary probabilities, in line with old expectations. We highlight the additional impact of both boundary conditions and integration contours on these issues.

hep-th↗

On the Bach tensor and quadratic curvature functionals

We study the critical points of a quadratic functional depending on the gradient of the Bach tensor on Riemannian four-manifolds, which generalize the Bach-flat condition. We show that, on every closed four-manifold, there exists a weak Bach-parallel metric, i.e. a critical point for this functional with respect to conformal variations: in particular, we prove that there exist infinitely many conformal classes which contain a unique minimizer for the functional, up to constant positive rescaling. Next, we analyze the global minima of the functional, i.e. metrics with parallel Bach tensor, relating these metrics to well-known variational problems. Using a version of de Rham's splitting theorem on complete four-manifolds, we provide a classification result for products of surfaces, exploiting the theory of conformal gradient solitons; we also construct a new explicit example of a Bach-flat metric which is neither locally conformally flat nor conformally Einstein and we characterize HCMU metrics on complete surfaces. Finally, we prove an equivalence between the Bach-parallel condition on 4D cylinders and the existence of critical metrics for a well-known quadratic curvature functional in dimension three: in this direction, we also prove a characterization of flat three-manifolds, under some curvature and finite energy assumptions.

math.DG↗

ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding

As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a systematic framework for unifying these behaviors. To bridge this gap, we introduce ParanoiaEval, the first benchmark for unified evaluation of risk-treatment capabilities in coding agents. Grounded in the well-established Avoidance-Transfer-Mitigation-Acceptance framework in software engineering risk management, ParanoiaEval operationalizes its 4 fundamental treatments for coding-agent settings and contains 200 evidence-controlled repository-level task pairs, each differing only in treatment-defining evidence. We further introduce dedicated metrics for risk-treatment violations and evidence responsiveness, using a human-calibrated agentic judge for reliable evaluation. Large-scale experiments on 8 representative models and a post-hoc human study reveal that (I) unnecessary risk treatment occurs in 11.2%-58.7% of runs despite explicit evidence, with substantial variation across agent configurations; (II) stronger task capability does not ensure more appropriate risk treatment, while treatment violations substantially harm developers' experience, establishing risk treatment as an independent capability dimension; and (III) agents exhibit systematic patterns consistent with established risk-management findings, suggesting that knowledge from human practice can guide the diagnosis and improvement of this capability.

cs.AI↗