Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Axial thermal expansion from free-energy minimization with temperature-dependent force constants in phonopy: a technical report, with $α$- and $ω$-Ti as the worked example

In a hexagonal crystal the $a$ and $c$ axes change their lengths at different rates with increasing temperature, and the thermal expansion is described by two coefficients, one for each axis. The equilibrium lattice parameters at each temperature are found by minimizing the Helmholtz free energy over the lattice parameters, and the coefficients are their logarithmic derivatives with respect to temperature. This report describes a procedure for this minimization as implemented in phonopy. The free energy is computed with harmonic force constants at a finite set of values of $a$ and $c$, and a free-energy surface over $a$ and $c$ is fitted to these values. In a variant, the harmonic force constants are replaced by temperature-dependent force constants from a self-consistent harmonic approximation, with the forces computed by a machine-learning potential fitted at each of these values. The $α$ and $ω$ phases of titanium are the worked example. Both calculations give negative thermal expansion of the $c$ axis of $α$-Ti at low temperature and none in $ω$-Ti, and they differ most for the $c$ axis of $α$-Ti. The axial coefficients are harder to compute than the volumetric one, because the two axes are coupled in the free energy. In both phases of titanium the thermal effect that lengthens $a$ also shortens $c$. The coefficient of the $c$ axis is then the sum of two terms of opposite sign that partly cancel, and a small error in either term gives a large relative error in the coefficient. The report measures how much each setting of the procedure changes the thermal expansion coefficients, with the $c$ axis of $α$-Ti as the most sensitive case.

cond-mat.mtrl-sci↗

Lifted Bellman Linear Programming for Offline Reinforcement Learning

Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets stabilized by target networks with exponential moving average (EMA) updates. Multi-step targets incorporate behavior-policy actions and therefore require off-policy correction. We instead impose in-sample Bellman optimality on the critic through inequality constraints. We formulate the Lifted Bellman Linear Program (LBLP), which lifts the linear programming characterization of Bellman optimality to the joint $(Q,V)$ space so that every constraint involves only state-action pairs in the dataset. Its unique minimizer is the in-sample optimal pair, and constraints along $K$-step segments of dataset trajectories leave this minimizer unchanged for any rollout policy and horizon. Under deterministic dynamics, this minimizer lies between the best dataset return and the optimal value. Relaxing the constraints into hinge penalties recovers the same solution above a finite penalty coefficient in the tabular case. Approximate Lifted Bellman Unconstrained Minimization (ALBUM) implements this relaxation with neural networks and detaches the $K$-step rollout targets by stop gradient. Its objective contains no squared regression onto bootstrapped targets, so it can be trained without target networks or EMA updates. Under deterministic dynamics, the LBLP solution is a stationary point of the detached update under a coefficient condition independent of $γ$ and $K$, and the inequality constraints allow discounted returns along dataset trajectories to serve as lower bounds without off-policy correction or action chunking. On OGBench, ALBUM uses a single critic with a Gaussian policy, matches the average performance of FQL, and is comparable to recent action-chunking methods, while using the fewest parameters and the least peak GPU memory among all compared methods.

cs.LG↗

Formal Verification of Proofs from Automated Theorem Provers for Higher-Order Logic

We identify common challenges and requirements for verifying proofs from automated theorem provers in the Dedukti logical framework and develop a general methodology for deriving encodings of calculus rules and proof steps, including clausification. We then apply this methodology to the EP calculus for higher-order logic and integrate it into the automated theorem prover Leo-III. The resulting prototype reconstructs about 80% of generated proof steps automatically, making Leo-III the first higher-order automated theorem prover to support independently checkable proof reconstruction and providing a basis for cross-system reuse. The implementation uncovered several bugs in Leo-III.

cs.LO↗

Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and Training

Large language models (LLMs) can improve solutions to verifiable scientific and algorithmic problems by spending additional computation at test time. Recent systems achieve strong results with increasingly elaborate evolutionary search harnesses or by updating model parameters during test-time training. We ask how much of this machinery is necessary. We introduce Hill Sampling, a form of hill-climbing optimization that repeatedly samples candidate programs from a frozen LLM, retains the best program found so far, and conditions all subsequent samples on that program. We evaluate the method on circle packing, sums and differences of sets, and Erdos' minimum-overlap problem using three open-weight models. Hill Sampling sets a new state of the art on circle packing among published methods, improves over the AlphaEvolve reference on Erdos' minimum-overlap problem, and achieves strong results on sums and differences of finite sets. The circle-packing and Erdos results require only hours of wall-clock time on eight NVIDIA H100 GPUs. We compare Hill Sampling against what is, to our knowledge, the largest application by parameter count of evolution strategies (ES) to LLM weights at test time. Surprisingly, when evaluating a method by the best program it generates, we find that learning model weights via ES is worse than invoking ES with a learning rate set to zero, i.e., using random weight-space perturbations to search for better models. Moreover, repeated sampling outperforms both ES methods, and Hill Sampling is the strongest of all. These results suggest a simple test-time compute allocation strategy: repeatedly sample edits to the best verified solution found so far before introducing additional complexity, such as adding archives, diversity mechanisms, evolutionary scaffolds, or test-time parameter learning.

cs.LG↗

Scalable Minimum-Volume Simplex Estimation with Non-asymptotic Analysis

We study the estimation of a $K$-dimensional simplex from $N$ i.i.d.\ points sampled uniformly from its interior; the observations are convex combinations of $K+1$ unknown prototypes. Existing polynomial-time estimators need cubic per-sample work or $O(NK)$ storage and are impractical at $N\sim 10^6$--$10^8$. We propose DeepMVSA, which re-expresses the minimum-volume principle in neural implicit form: a lightweight coordinate network generates the mixing weights and a triangular LU-type parameterization the dual simplex matrix, reducing the trainable-state memory to $O(K^2)$, independent of $N$, and the cost per data pass to $O(NK^2)$. We prove a non-asymptotic sample-complexity bound of the polynomial-time benchmark order for a localized surrogate estimator; an oracle inequality for every global minimizer of the neural objective, with volume-inflation control and an explicit shrinkage bias; a conditional end-to-end error budget separating statistical, approximation, optimization, and enclosure-residual terms on an explicit envelope event; and two-point lower bounds: at any noise level $σ>0$ fixed independently of $N$, the $N^{-1/2}$ scaling is unimprovable in its $N$-exponent. Experiments with up to $N=10^8$ synthetic observations are consistent with the predicted accuracy and scaling, and feasibility on real scenes of $\sim 10^7$ pixels is demonstrated.

stat.ML↗

Exponential Quantum Advantage in Testing Fourier Dimensionality

A boolean function $f$ has Fourier dimension $k$ if its nonzero Fourier coefficients span a subspace of dimension $k$. We consider the property testing task of determining whether a function has Fourier dimension at most $k$, or is $ε$-far from being so. We show that there is a $O(k/ε)$-query quantum property tester for this problem, which we show to be almost optimal. Combined with Gopalan et al.'s classical lower bound of $Ω(2^{k/2})$, this demonstrates an exponential quantum advantage for this task. We complement this result with a $\tilde{O}(2^{k/2}/ε)$ classical tester, giving a quadratic improvement over the previous best tester, and essentially settling the classical query complexity.

quant-ph↗

Exact Second-Order Zarankiewicz Numbers for Complete-Graph Incidence Families

For $n\ge6$, $m=\binom n2$, let the complete-graph incidence family on $K_n$ have the vertices of $K_n$ as columns, its edges as rows, and the incidence graph as one-edge graph. The universal cell bound of Löfberg and Qi gives $z_2(m,n)\le Z(n):=\lfloor n(n-1)(n+2)/4\rfloor$. We implement the nested one-factorization construction of that family and determine exactly what it certifies. For $n=2q$ with $q$ an odd prime, $q\ge5$, the construction has no hole and the cross-factor transfer equations apply verbatim, giving $z_2=z_{SL}=z_{RL}=Z(n)$; $n=6$ is settled by a separate cyclic witness. For odd $n=2p+1$ the near-perfect one-factorization yields Hamilton paths closed into odd cycles, and the transfer argument breaks: that scheme certifies only $R(G_p)\le Z(n)$. The orders $n=7,8,9,12,13,16,17,18,20,21$ are settled by explicit configurations of a different shape, each attaining the cell bound and satisfying $(\mathrm{RW}3^+)$, obtained from a larger configuration by deleting vertex stars and repairing the restricted grid. At each of these orders $z_2=z_{SL}=z_{RL}=Z(n)$; all remaining orders are conjectural. The first settled order $n=7$ is the only one whose grid has a hole, grounded by a zero-companion rule. Machine-readable configurations and a certificate checker accompany the paper.

math.CO↗

The Uncontrolled Variable: Vision-Language Refusal Is Conditioned on the Image-Attachment Interface, and Not Robust to Irrelevant Image Properties

We show that aligned vision-language models also condition refusal on a property of a request's form: whether an image is attached, holding everything the request asks fixed. Attaching a blank canvas, an image that cannot be read, cannot relate to the request, and is byte-identical across every prompt in its condition, shifts benign refusal by tens of points. The shift is not blanket caution but a threshold shift: genuinely neutral instructions are almost unaffected (<=2 percentage points on three of four hosted models) while borderline-benign prompts move +23 to +51 points, so the cost falls on sensitivity-adjacent traffic, meaning benign questions about privacy, self-harm, violence and illegal activity. There is a benign reading of such a threshold, namely that attachment correlates with risk in real traffic, and we take it seriously; a black-box study cannot measure that correlation and we do not claim to. What it can test is whether the response to attachment is robust to variation carrying no information about the request, and on four independent measurements it is not. It varies with canvas colour and pixel count. Its sign inverts across checkpoints. It survives an explicit instruction to disregard the image. And on one model it fires on a bare assertion that an attachment exists, with nothing attached and the modality word contributing none of it. Finally we price it. On a matched harmful set the same canvas does lower attack success, so the cue buys something. But the charge is decoupled from the purchase: the checkpoint with the least harmful headroom we measure, completing only 2% of plain harmful requests, still pays the benign cost in full, and across our models the harmful-side denominator falls as alignment improves while the benign cost does not track it down. Image presence is not a conservative default that a deployer chose and priced. It is an uncontrolled variable.

cs.CR↗

EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics

We study coding agents for long-horizon, dexterous robotics and ask whether their solutions can provide scalable supervision for learning general robot policies. To test this, we develop EMBODIEDSWE-BENCH, a simulation benchmark for coding agents spanning contact-rich manipulation, deformable objects, and long-horizon tasks requiring up to half an hour of continuous interaction. We find that frontier coding agents can solve complex long-horizon tasks and transfer prior solutions across both tasks and embodiments. We also design supporting tools that help agents more effectively solve these tasks. However, the resulting solutions require substantial iterative interaction and are typically specialized to individual task instances. We therefore introduce EMBODIEDSWE-GEN, which expands a single solution from coding agent into large diverse trajectories for training a VLA. VLA performance improves with more generated demonstrations, and agent-aided diversification improves generalization to held-out task variations. We also show that a VLA finetuned solely on coding-agent-generated simulation demonstrations completes a long-horizon task on real robot. Together, our framework uses coding agents to solve complex robotics tasks and turn verified solutions into scalable supervision for robot policies.

cs.RO↗

NV-Reason-CT: 3D Visual Language Model for CT Analysis

We present NV-Reason-CT, a generative vision--language model for chest and abdominal CT combining native 3D visual encoding with radiologist-guided reasoning. The model couples a native 3D vision transformer with a language model, passing all visual tokens and their explicit 3D coordinates into language decoding without further spatial token merging. This retains volumetric spatial information within the vision encoder and through the language model's positional encoding during joint processing with text. We train on a curated corpus of approximately 550,000 multimodal instruction examples from 70,111 unique CT image inputs, combining standardized reports, abnormality-focused and anatomy-specific questions, multi-turn interactions, and radiologist-authored reasoning from recorded and transcribed expert CT interpretations. Expert annotations provide direct supervision and guide additional report-grounded synthetic reasoning. End-to-end supervised fine-tuning (SFT) is followed by Group Relative Policy Optimization (GRPO), with verifiable rewards over chest and abdominal abnormality sets. The model supports abnormality classification, report generation, and interactive reasoning with reviewable observations, differential diagnoses, and uncertainty. Evaluation spans public CT benchmarks and a held-out NIH cohort. On CT-RATE, NV-Reason-CT achieves a macro-F1 of 0.614 and macro-AUROC of 0.871 without a task-specific classification head; generated reports achieve a report-derived macro-F1 of 0.592. In a preliminary study with expert radiologists, AI-assisted review received favorable confidence ratings and was associated with a 50% reduction in average reported interpretation and reporting time. We release the model and training code to support reproducible research on explainable AI for volumetric medical imaging.

cs.CV↗

Refuting the QAOA fixed-angle conjecture

The fixed-angle conjecture for the quantum approximate optimisation algorithm (QAOA) is shown to be false for depth-2 QAOA on 9-regular graphs, and proven to be correct for depth-1 on any regular graph, as well as for any depth on all 2-regular graphs.

quant-ph↗

Transferable Evidence Reconstruction for Longitudinal Glucose Representations

Long physiological recordings contain many routine measurements, while predictive information often lies in rare events, sustained burden, and recurring patterns. These properties can be computed as label-free evidence, but directly using them as features leaves limited labeled data to separate reproducible associations from sample-specific ones. Learning to reconstruct evidence can exploit unlabeled recordings, yet joint reconstruction does not explicitly require the decoding rule to transfer across individuals. We introduce transferable evidence reconstruction (TER): a Ridge regressor fits evidence from representations in one group and predicts it in an identity-disjoint group without refitting. The transfer error trains the encoder through the differentiable fit. For continuous glucose monitoring (CGM), clock-aware encoding preserves the multi-day content and timing needed for evidence recovery. Matched interventions connect the gains to reduced fitting-group sensitivity, with structured targets improving on raw recovery. Across ten leading CGM and time-series baselines, TER sets a new best metric on 12/14 phenotype tasks and exceeds the strongest prior overall PR-AUC/ROC-AUC/Macro-F1 by 4.95/4.43/0.66 percentage points; the PR-AUC and ROC-AUC gains are $2.6\times$ and $2.2\times$ the respective gaps between the two strongest baselines. Meal-response and future-CGM studies further demonstrate predictive utility. TER thus uses meaningful signal properties to supervise not only what a representation preserves, but how reliably it can be read across individuals.

cs.LG↗

AI in Science: Early Insights

Scientific progress is a key driver of economic growth and prosperity. There is great excitement - but also concerns - about the impacts of AI on science, but so far little data. We provide early insights on this from three data sources: a sample of 15 million Gemini interactions, an inventory of over 2,600 specialized AI models across disciplines, and a survey of over 600 scientists. We map these data to a new taxonomy of scientific tasks to study how scientists are using AI. Four main findings emerge. First, we find broad adoption and coverage: scientists use AI more than most other occupations. Specialized AI models have broad disciplinary coverage and are highly cited. Nearly half of the scientists surveyed report using some form of AI every day. Second, we document evidence that LLMs (proxied through Gemini usage) and specialized models act as complements-- LLMs are used for general analysis, coding, and manuscript preparation, while specialized models provide domain-specific predictions, data generation and classification. Third, scientists report large productivity gains from using AI: a saving of nearly 7 hours per week, time which is primarily re-invested in more research. Finally, we show that AI is already changing the scientific process. As some stages of scientific research become easier, bottlenecks shift downstream. Scientists report an increased backlog of untested hypotheses and substantial demand for output verification. Our findings suggest that AI holds significant potential to increase scientific productivity. However, as with other sectors, its ultimate impact will be governed by complex task interdependencies and investment into the elimination of emerging bottlenecks.

econ.GN↗

The Price of Removing the Scaffolding: Maxwell wrote something about Ampère that years later would be done to him

In the history of physics, the name scaffolding has been given to whatever served to raise a theory and is afterwards removed because, it is said, it is no longer needed. This article documents, from the originals, what it cost Maxwell's theory when that method was applied to it: what was removed from it, who did it, with what argument, and how over the years it had to be recovered because it was understood that it was needed after all. It also records what Maxwell himself missed in Ampère, what he calls the traces of the scaffolding: the template with which Ampère himself had given shape to his work.

physics.hist-ph↗

Wiman-Valiron inequalities in the unit disk outside sets of finite logarithmic measure

We give affirmative answers to both parts of Question 2.6 posed by Grosse-Erdmann (2025) concerning Wiman--Valiron inequalities in the unit disk. For every unbounded analytic function in the disk, we establish the proposed iterated-logarithm inequalities outside exceptional sets of finite logarithmic measure. The corresponding power estimate is a corollary. Both conclusions follow from a variance bound for Khinchin families and a classical estimate for their largest atom. The multiplicative constants in the main inequalities can be chosen absolute. We also obtain a disk analogue of Rosenbloom's composition estimate, with an explicit boundary prefactor. The key step combines a boundary change of variable with a monotone auxiliary function whose derivative is exactly the variance of a rescaled member of the Khinchin family. A classical example shows that the leading logarithmic exponent $1/2$ cannot be decreased.

math.CV↗

Unrolling Lanczos for Ideal Low-pass Graph Filter Approximation

Low-pass (LP) filtering is a fundamental operation in graph signal processing (GSP). Among finite-order nodal-domain methods, Lanczos-based filtering provides more accurate approximations of ideal LP filters than Chebyshev polynomial methods. We show that the approximation of Lanczos filtering can be further improved through algorithm unrolling and data-driven parameter learning. The key insight is that, because ideal LP filtering is a projection operation into the low-frequency eigen-subspace $\cS_K$, instead of approximating individual eigen-pairs of a graph Laplacian $Ł$ as done in classical Lanczos, an unrolled Lanczos network can directly approximate $\cS_K$. Specifically, we first establish a theorem identifying properties of the Lanczos tridiagonal matrix $\T_m$ that promote accurate approximation of the low-frequency eigen-subspace $\cS_K$. Guided by this theory, we relax the orthogonality constraint on Lanczos vectors, resulting in Ritz vectors that better span $\cS_K$. To ensure numerical stability, we constrain $\T_m$ to be similar to a symmetric matrix, thereby guaranteeing real-valued eigenvalues. Experimental results on random graphs and learned graphs in two natural language processing (NLP) tasks show that our unrolled Lanczos network achieves superior ideal LP filter approximation compared to the classical Lanczos method.

eess.SP↗

Information geometry of emergent symmetry quotients and their weak unfoldings

A regular observed statistical model may converge to a limit in which a previously identifiable signed parameter becomes identifiable only modulo a reflection. We study the local information geometry of this transition. For a twice differentiable Hellinger embedding with an exact limiting reflection, the observed displacement is forced into the two-jet form \(\varepsilonλJ_-+λ^2J_+/2\), up to higher-order terms. The mixed jet restores the sign away from the symmetric face, whereas the even jet is the first tangent inherited by the quotient. After nuisance elimination, a positive Gram determinant yields a nondegenerate cross-cap two-jet. The associated local asymptotic theory has three regimes governed by \(τ_n=\sqrt n\,\varepsilon_n^2\): regular signed LAN, a critical curved Gaussian subexperiment, and a quotient regime with the \(n^{-1/4}\) signed scale. We prove that the same parabolic critical experiment persists for predictive likelihoods along a single stationary dependent trajectory. The limiting quotient has a regular Fisher metric in the invariant coordinate, while its pullback degenerates in the signed coordinate. For a solvable CIR--OU benchmark motivated by coherent sea-clutter observations, we derive the quotient Fisher metric and curvature explicitly and show that the curvature is strictly negative. We also determine the restricted holonomy of the full Amari family: \(\operatorname{Hol}_0(\nabla^{(a)})=SO(2)\) for \(a=0\), whereas \(\operatorname{Hol}_0(\nabla^{(a)})=GL^+(2,\mathbb R)\) for \(a\neq0\). The results separate the intrinsic geometry of the limiting quotient from the transverse geometry of its weak unfolding.

math.ST↗

EgoSpeedUp: Transferring Human Manipulation Tempo to Robot Policies

Robot manipulation policies trained through imitation learning inherit not only the demonstrated behavior but also the conservative execution tempo of robot demonstrations. Existing acceleration approaches can execute faster than the original demonstrations, but determine the appropriate acceleration primarily from robot-side information or a predefined set of tempo factors, leaving open how to obtain a task-appropriate reference for how fast each manipulation phase should progress. We introduce EgoSpeedUp, a framework that uses human manipulation as temporal supervision for robot imitation learning. Our key insight is that human demonstrations naturally reveal task-appropriate, phase-wise manipulation tempo. Given slow robot demonstrations and human demonstrations of the same task, EgoSpeedUp aligns corresponding manipulation phases, estimates their relative execution tempos from multiple human demonstrations, and transfers the resulting phase-wise tempo by retiming the robot demonstrations. The retimed demonstrations are then used for standard behavior cloning, allowing the robot to retain its executable manipulation behavior while learning to perform it at a human-informed tempo. Across two real-world manipulation tasks, EgoSpeedUp improves the task success rate by an average of 25 percentage points (pp) while reducing successful execution time by 36.5%. These results demonstrate that human manipulation tempo provides an effective temporal reference for learning faster and more reliable robot policies.

cs.RO↗