Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Structure of irreducible homomorphisms to/from free and injective modules

Let $R$ be a commutative noetherian local ring. We extend the work of the author and Takahashi by investigating the structure of irreducible monomorphisms originating from free modules in the category of finitely generated $R$-modules. In the case where $R$ is complete, we further study the structure of irreducible monomorphisms and epimorphisms to and from injective modules in the category of artinian $R$-modules.

math.AC↗

Long-time Korteweg-de Vries approximation for the Fermi-Pasta-Ulam-Tsingou system

We prove that, as the lattice spacing $h$ tends to zero, general solutions to the infinite Fermi--Pasta--Ulam--Tsingou (FPUT) system can be approximated in $L^2$ by two counter-propagating Korteweg--de Vries (KdV) waves on time intervals of order $\log(1/h)$. This resolves an open question raised by the first author and collaborators~\cite{HKY2021}. Our proof combines the FPUT conservation law with an $h$-uniform local well-posedness theory in $L^2$ and persistence of Sobolev regularity. We also introduce a frequency-localized auxiliary equation to overcome the difficulty of comparing the Fourier restriction norms associated with the FPUT and KdV flows.

math.AP↗

FRAM: Trajectory-Guided Visual Feature Selection for Compact Language-Conditioned Robot Manipulation

Vision-Language-Action models achieve strong performance in robot manipulation, but often require large numbers of parameters. In this work, we propose the Future Representation Action Model (FRAM), a small policy that explicitly links the future end-effector trajectory to the current visual input. FRAM uses the image coordinates of the predicted trajectory as spatial pointers and reads local visual features related to the motion from the current image. This organizes the information for action generation into the reference position (Where), the visual state (What), and the future motion (Future). Trajectory labels are generated automatically from demonstrations and camera geometry, so no manual annotation is needed. With 138.7M parameters, including a frozen language encoder, FRAM reaches an average success rate of 92.2% over the four standard LIBERO suites, close to the 94.2% of $π_0$ with 3.3B parameters. Without extra training, it also reaches an average of 67.3% on LIBERO-Plus. Ablations confirm that both the future trajectory and the local visual features improve performance and robustness. On a real dual-arm UR5e, FRAM stacks cups using only wrist cameras, including choosing and switching between the left and right arms. These results show that selecting visual information based on future motion is an effective way to obtain both high performance and robustness in a small robot policy.

cs.RO↗

Group recovery after trimming, and level-free flagging, in robust clusterwise regression

Trimming methods for robust clusterwise regression discard a fixed fraction of the data. Too low a level breaks the fit; too generous a level can trim away a small group. We first propose a group-recovery step that can follow any trimming or flagging method: it searches the discarded units for a line, tests whether those near it form a peak rather than a band, and restores the line as a group when the likelihood of Gaussian groups plus uniform noise improves by a margin like that of the Bayesian information criterion. In simulations it repaired the failures of a generous TCLUST-REG level with unequal groups, but not with three or four groups. The second proposal, ESF (exact-subsample flagging), is a flagging procedure without a trimming level: it solves the clusterwise least-squares problem exactly on small subsamples, flags units far from the best fit, and draws later subsamples from the rest. Two constants stand in for the level: a subsample size, set from a lower bound on the smallest group proportion, and a cap on the flagged set. It is meant for data about whose contamination nothing is known: TCLUST-REG at the fixed level 0.30, followed by reweighting and the recovery step, was as accurate as ESF on average up to a fifth of outliers, and higher levels were more accurate beyond. The flagged fraction estimates the contamination only when the errors are close to Gaussian. On taxi fares with known tariffs, ESF found both in every sample.

stat.CO↗

Low-Order Refined Preconditioning for Spectral/hp Element Method for Complex, 3D Geometries

Low-order refined (LOR) preconditioning replaces a high-order operator with a low-order discretisation on a refined nodal mesh. For tensor-product elements, the two operators are spectrally equivalent with bounds independent of the polynomial order $P$, but the construction does not extend directly to simplex and mixed-element discretisations. This work makes two contributions: it extends LOR preconditioning to simplex and mixed-element discretisations, including triangular, tetrahedral, and prismatic elements, and establishes a generalised Vandermonde transformation linking the modal and nodal LOR formulations, showing that the resulting preconditioned spectra and Krylov convergence are independent of the high-order basis. Numerical experiments show controlled iteration growth on triangular meshes despite increasing condition number, and controlled iteration counts up to $P=5$ on tetrahedral, prismatic, and mixed-element meshes. A single algebraic multigrid V-cycle per outer iteration gives the best balance of iteration count and cost. The method is applied to a production incompressible Navier-Stokes simulation of a race-car front-wing and wheel configuration, discretised on a mesh of $2.87\times10^6$ mixed prismatic and tetrahedral elements giving $32.2\times10^6$ pressure degrees of freedom at $P=3$. LOR reduces the mean pressure conjugate gradient (CG) iteration count from $235.3$ to $5.5$, and the pressure-solve time over 1000 timesteps by 16.1%, relative to the default production static-condensation diagonal preconditioner in Nektar++. The one-time cost of constructing the LOR preconditioner is amortised over the production simulation.

math.NA↗

DOHF: Online Diffusion Fine-tuning with Doob's $h$-transform Guidance

Reward-based diffusion fine-tuning faces practical challenges when desirable outcomes are rare or conditioning corrections are costly to estimate. In this work, we propose Diffusion Online $h$-guidance Fine-tuning (DOHF), which turns Doob's $h$-transform into a practical online training algorithm. DOHF assigns optimality weights to generated samples, estimates the normalized local correction $\nabla\log h$ under the current rollout policy, and distills it directly into the generative model. Theoretically, we characterize the population-optimal DiffusionNFT update as well as the various classfier free guidance methods through a unified $h$-transform perspective. Methodologically, our framework accommodates black-box and non-differentiable rewards without additional network evaluations. We further show improved alignments under three empirical scenarios. Our work demonstrates how adapting probabilistic conditioning through inexpensive estimation and iterative distillation can improve generative learning across statistical sampling and visual generation.

cs.LG↗

Devol-ONE: One Autoregressive Mixture of Transformers to Unify Vision-Language-Action and Latent World Modeling

Vision Language Action (VLA) models condition actions directly on current visual and language context, without an explicit account of how the scene evolves under candidate actions. World Action Models (WAM) attempt to address this limitation by predicting future states, but existing designs keep prediction and policy learning architecturally separate, connecting them only through the predicted output, whether through pixel space video generation or a latent forecasting module trained independently of the policy. We present Devol-ONE, a Mixture of Transformers architecture that unifies vision language understanding, latent world dynamics prediction, and action generation within a single autoregressive framework. Instead of encoding vision language tokens once and feeding them to the action expert, Devol-ONE runs autoregressive prediction jointly across a vision language stream and a V-JEPA pretrained dynamics stream, attending to the vision language key-value cache at every layer to forecast future latent states under language guidance. The action expert is in turn shaped continuously by semantic reasoning and predicted physical dynamics rather than by a fixed representation computed in advance. Extensive experiments are conducted on LIBERO, LIBERO-PLUS, RoboTwin2.0 along with real-world evaluation on Flexiv single-arm and dual-arm setups. Ablation studies show the effectiveness of dynamic stream prediction and layer-wise unified attention to validate our model architectural coherency.

cs.CV↗

Assessing the impact of climate change and rising temperature on life insurance portfolios

Climate change may materially affect long-term life insurance liabilities by altering both the level and seasonal pattern of mortality. This article develops a multi-population mortality framework that combines a Hermite spline model with a distributed lag non-linear model to capture age-, region-, and temperature-specific mortality effects. We apply the framework to mortality and temperature data from 15 Spanish NUTS-2 regions and project future mortality under three shared socioeconomic pathway (SSP) scenarios. We then assess the implications for a hypothetical whole life insurance portfolio through expected death-benefit payments and portfolio profit and loss. The results reveal an important seasonal offset: warmer conditions reduce expected payoffs during winter periods but increase them during summer periods, with these effects becoming more pronounced under more severe climate scenarios and for policies issued in later years. Over longer horizons, adverse summer mortality effects become increasingly important. The portfolio analysis further shows that climate-related mortality risk can materially increase the dispersion and downside risk of portfolio outcomes, particularly under SSP5-8.5. These findings highlight the importance of incorporating temperature-related mortality effects into long-term life insurance liability projections and risk assessment.

stat.AP↗

Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction

Visual token reduction is an effective way to accelerate multimodal large language models (MLLMs), but performance deteriorates rapidly under extremely low token budgets. Existing work has explored both visual-token selection and training-based adaptation to reduced visual inputs. We take a step further by asking how a heavily compressed MLLM should learn from the states induced by its own generations. This setting naturally calls for on-policy self-distillation: a heavily compressed model is supervised on the states induced by its own generations, while its full-token counterpart serves as an information-rich teacher. Based on this insight, we propose LT-OPD, a training framework for extreme visual-token reduction. The student rolls out responses with only a small fraction of visual tokens, and a frozen full-token copy of the same MLLM provides distributional supervision along these student-generated trajectories. To stabilize on-policy learning when visual evidence is severely limited, we further introduce a budget-level curriculum that progressively decreases the token budget during training. Across nine benchmarks on Qwen3.5-4B, LT-OPD raises average retained performance under 5% visual-token retention from 68.6% to 82.3%, outperforming training-free, training-based, and reinforcement-learning baselines at the same budget. The gains transfer consistently to Qwen3.5-9B, GLM-4.6V-9B, and LLaVA-OV-1.5-4B. LT-OPD also reduces KV-cache usage by 85.2% and prefill FLOPs by 85.4% without additional inference overhead, demonstrating that on-policy learning can substantially recover capabilities lost to extreme visual-token reduction.

cs.CV↗

Convergent Plug-and-Play Image Restoration with Annealed Noise Levels

Plug-and-Play (PnP) methods solve imaging inverse problems by incorporating deep denoisers into iterative optimization algorithms. Although practical implementations often decrease the denoiser noise level $σ$ along iterations, most existing convergence analyses assume a fixed denoiser. In this work, we establish convergence guarantees for a broad family of Plug-and-Play algorithms with annealed noise level, spanning deterministic methods (RED--GD and PnP--PGD) and stochastic methods (SNORE, equivariant RED, and a variant of PnP--Flow). For each method, we identify an explicit, nonconvex objective associated with the terminal denoising level and prove asymptotic stationarity of the iterates with respect to this objective. Our analysis does not prescribe any decay rate for the noise schedule, and our assumptions cover both learned gradient-step denoisers and exact MMSE denoisers. Overall, our theoretical results bridge the gap between existing PnP convergence theory and the decreasing-denoising practices used by state-of-the-art image restoration methods. We empirically demonstrate the benefits of such schedules and illustrate the predicted convergence behavior on several imaging inverse problems, including inpainting, super-resolution, demosaicing and tomography.

math.OC↗

Spectral Reality for Certain Fourth-Order Nonsymmetric Finite-Difference Laplacians

Finite-difference discretizations of Laplace operators yield discrete Laplacians whose spectral properties are closely tied to the stability, convergence, and physical fidelity of numerical solvers. While symmetric discretizations are well understood, many high-order finite-difference schemes produce nonsymmetric matrices for which rigorous spectral analysis is overlooked. Surprisingly, we prove that two fourth-order schemes with Dirichlet boundary conditions yield nonsymmetric discrete Laplacians whose spectra are purely real and strictly positive. Computer-assisted computations show that this property does not hold for higher-order extensions of these schemes in general.

math.NA↗

Not Too Hard, Not Too Easy: Learning from Intermediate States for LLM Structured Reasoning

A common principle of effective learning is to practice material that is neither already mastered nor too difficult to permit progress. We ask how to apply this principle to structured reasoning tasks such as Sudoku and maze solving. In these tasks, a model can repeatedly revise an incomplete or incorrect candidate solution until it satisfies the problem's constraints. The intermediate candidate solutions along this trajectory provide natural training examples: some are already solved, some cannot yet be repaired by the model, and others lie at its current frontier of achievable progress. We therefore investigate whether pretrained language models can learn to revise such states and whether training on states at this frontier improves reasoning more broadly. To achieve this, we couple a pretrained language-model backbone with a recurrent updater that repeatedly revises an explicit solution state, using the same parameters at every update step. We further introduce Frontier-Oriented Curation Using Self-trajectories (FOCUS), which selects training states from trajectories generated by the current model. FOCUS measures how much the model improves each state within a fixed number of recurrent updates and prioritizes states from which it can make substantial progress. With Qwen3-1.7B, FOCUS achieves 64.4% exact solve accuracy on Sudoku-Extreme and 91.1% on Maze-Hard, with similar gains observed across five Qwen and Llama backbones spanning 1.7B to 8B parameters. We further observe zero-shot transfer in the adapted LLM to mathematical reasoning and code execution, even when the recurrent updater is disabled and no downstream fine-tuning is performed.

cs.AI↗

What Does a Skill Actually Do? Estimands and Evaluation Validity for Tool and Skill Use in LLM Agents: A Critical Review

Reported improvements from tools and reusable skills in large language model agents refer to different comparisons. This critical narrative review examines what these evaluations estimate and which conclusions their designs support. The review checks the roles of one hundred cited papers and extracts focal evaluation designs in detail from thirty-five studies. Targeted readings of thirty-five additional published or accepted studies broaden coverage of tool creation, memory, interactive benchmarks, reliability, and risk. Designs are characterized by treatment contrast, target population, outcome, budget constraint, summary measure, and identification assumptions. Analytic decompositions and counterexamples show that pairing runs on the same task does not itself identify an invocation effect when evaluation conditions on a trigger within the treated run. Paired gain and regression counts describe discordance under the coupling protocol rather than the share of tasks whose expected outcomes worsen. Total effects of deploying a module answer a different question from efficiency under a common budget. Comparisons across studies distinguish curated skill provision from retriever replacement, task populations from triggered subsets, and preparation costs from marginal usage costs. Publication status and reading depth are recorded. The review provides a methodological synthesis and a reporting checklist to help align claims about tools and skills with the comparisons their evaluation designs support.

cs.SE↗

FloodDiffusion 2: Efficient and Path Controllable Streaming Motion Generation

We present FloodDiffusion 2 (FD2), an efficient and controllable framework that builds upon FloodDiffusion (FD1), a state-of-the-art streaming motion generation model. While FD1 produces plausible motion, it suffers from low efficiency and limited controllability, as its attention design requires repeated computation over the entire history, and it lacks precise trajectory control for real-world applications. To address these limitations and improve generation quality, FD2 introduces three advances. First, Partial Attention makes finalized history representations independent of the active window, enabling KV-cached inference and shared-history packing for efficient training. Second, we establish a necessary-and-sufficient Bregman criterion for regression losses to preserve diffusion's conditional-mean velocity field. This criterion guides an FK-induced quadratic loss that incorporates motion geometry without online FK evaluation. Third, FD2 introduces precise path conditioning to control the character's root trajectory while preserving natural body motion. Experiments show that FD2 reduces training computation by 4.6$\times$ and accelerates denoising by 11.29$\times$, reaching 2.303 ms per update on long sequences. Alongside these efficiency gains, FD2 improves motion quality over FD1 and achieves state-of-the-art FID scores among streaming methods, with 0.048 on SEED and 0.053 on HumanML3D.

cs.CV↗

Learning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product Markets

Autonomous large language model (LLM) agents operating in multi-product markets must make sequential decisions under information asymmetry and resource constraints. We develop a machine learning approach for training such agents to act effectively as sellers in a multi-item bargaining environment, where a seller concurrently negotiates a catalog of substitutable assets across a pool of independent buyers. Buyers hold private, heterogeneous valuations across products, and each can purchase at most one item. Facing limits on total communication turns, the seller must dynamically match buyers with the most profitable products considering their private valuations, while strategically allocating its limited interaction budget toward combinations of greater potential value. We formalize this problem as a Partially Observable Markov Decision Process using a structured, four-part message protocol that maps natural language into a parsable and regulated decision space. Using this formalization, we design a post-training method using Reinforcement Learning from Verifiable Rewards (RLVR). To evaluate this framework, we construct a multidimensional metric suite that quantifies constraint adherence, seller surplus extraction, and allocation quality. Our trained seller agent learns to match limited inventory to buyers more effectively, matching or outperforming trillion-parameter frontier models in both seller surplus extraction and buyer-product allocation quality. Finally, these learned strategies generalize robustly to unseen market structures, correlated valuation distributions, and price ranges not encountered during training.

cs.AI↗

An EPTAS for Vector Scheduling with Time Intervals

We study vector scheduling in which each job is active during a fixed time interval. A job uses several resources and stays on one machine for its entire interval; its resource requirements may depend on the machine. The objective is to minimize the largest resource load over all machines and times. For $r$ machines and $d$ resources, we give a deterministic $(1+\varepsilon)$-approximation in $f(r,d,1/\varepsilon)N^{O(1)}$ time, where $N$ is the binary input length. This gives an efficient polynomial-time approximation scheme for fixed $r$ and $d$, extending approximation schemes for scalar temporary tasks assignment. The algorithm merges jobs into blocks whose time intervals are fixed before any machine is chosen, and assigns the blocks by dynamic programming over a balanced recursive split of the time line. We also prove strong NP-hardness and an exponential lower bound in $1/\varepsilon$ under the Exponential Time Hypothesis, already for two identical machines and one resource.

cs.DS↗

Neuro-Symbolic Indirect-Call Analysis under Opaque Pointers

Resolving indirect calls is central to call-graph construction for C. Scalable type-based analyses such as MLTA use type information in LLVM IR to associate indirect calls with functions assigned to the corresponding structure fields. However, a single pointee type often misrepresents the memory a pointer addresses, and LLVM 17 removed pointee types in favor of opaque pointers. Therefore, field-sensitive analyses lose their matching key. Recovering the erased types restores the matching key but still misses the relation that the type encoded: which functions the program assigns to the field. We present Facet, to our knowledge the first analysis that reconstructs this dispatch relation over opaque IR. Facet identifies the structure field from which an indirect call loads its function pointer. It separately recovers the functions assigned to that field through initializers, stores, and aggregate copies. It then joins the two by field identity, without requiring an end-to-end value-flow path. Facet classifies proposed call-graph changes under distinct evidence rules for edge addition and removal and records the assumption behind each refinement. An LLM decides only the residual cases among symbolically bounded candidates. One analysis yields both a recall-preserving call graph and a refined call graph. On 14 C programs, Facet reduces the mean target-set size from 25.9 to 5.2 and raises observed recall from 0.79 to 0.99. Its recovered field identities agree with typed IR at 98.1% of jointly resolved sites. Applied to bug detection, the refined call graph found 17 deep bugs in C software from nginx to the Linux kernel, three of them latent for over a decade; 12 are confirmed.

cs.SE↗

Yorùbá in Unicode: An Overview of a Problem

There is a recurrent problem in the writing of Yorùbá on the internet and on the computer that has proven intractable over the years. The language, along with other African languages that depend on diacritics for disambiguation, requires a small set of precomposed characters that Unicode does not encode. This has forced writers and digital systems to rely on combining character sequences that behave inconsistently across platforms, corrupt under font substitution, and fail in search. This paper documents that failure across a range of real world contexts, from published books to web platforms to mobile keyboards, using personal and empirical evidence. It identifies Unicode's NFC normalization stability policy as the structural constraint that prevents a straightforward fix, arguing for direct intervention of the Consortium in solving the active problem, proposing a formal encoding request for the four core Yorùbá characters as the most durable path to resolution.

cs.CL↗