Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 667 records · Page 37Linked to original sources

On the Tightness and Computational Tractability of Higher-Dimensional Confidence Sequences

Modern sequential monitoring problems often involve multiple metrics, where we monitor several data streams simultaneously and may act once the evidence is strong enough. Confidence sequences (CSs) are a natural tool for such continuous monitoring. However, for bounded vector means, existing multivariate CSs are either tight but computationally intractable, or fast to compute but conservative. To address this, we study three lifts of one-dimensional betting-based CSs to higher dimensions: a weighted Bonferroni region, an equivalent max-wealth form, and a portfolio region. The portfolio is typically much tighter, especially in higher dimensions, but its boundary and properties such as volume are not available in closed form. To make this tighter construction usable, we propose tractable outer approximations of the portfolio region that preserve statistical validity: a bounding box, an $\ell_p$-ellipsoid, and their intersection. We prove set relations among all constructions and show empirically that these approximations (i) achieve regions close to the intractable portfolio, (ii) substantially outperform existing tractable multivariate CSs, and (iii) enable practical use cases such as multi-metric A/B testing and model comparison.

stat.ML↗

The Riemann-Siegel remainder as a fractional summand

Starting from the Riemann-Siegel decomposition of $ζ(s)$ given in Siegel's 1932 paper, we introduce a change of variable, $t=I(T)$, that replaces Siegel's coupled pair "imaginary part $t$ / summation index $m$" with one real index $T$. We then split Siegel's single remainder integral $R$ into two exact pieces, $R_{1ps}$ and $R_{2ps}$, and prove that $R=R_{1ps}+R_{2ps}$, so that $ζ=Σ_1+R_{1ps}+Σ_2+R_{2ps}$. Our central observation is that each of these remainders is nothing more than one additional, fractional partial summand appended to its Dirichlet sum: $ζ(s)=\sum_{n=1}^{m}n^{-s}+\hat{d}_1(m+1)^{-s}+χ(s)\sum_{n=1}^{m}n^{s-1}+\hat{d}_2χ(s)(m+1)^{s-1}$, with $\hat{d}_1,\hat{d}_2$ real numbers (always positive on the critical line), the fractions of those two summands that are used. As a corollary, when $σ=\frac{1}{2}$ one has $d_1=d_2$ (equivalently $\hat{d}_1=\hat{d}_2$); this fact is formally verified in Lean. Also, with this rescaling the remainder terms are nearly periodic in $T$ with period one, converging to a fixed waveform in the fractional part of $T$. This decomposition of Siegel's $R$ was discovered through experimental mathematics using a spiral visualization of the partial sums, described later in the paper. We also discuss a number of other observations, including what we call the yin yang curves, the zero counting function, and ovals of equal length leg loci.

math.NT↗

On injective dimension of the conormal module

Let $Q$ be a Gorenstein local ring, let $I$ be a proper ideal of $Q$, and set $R=Q/I$. We investigate when finite injective dimension of the conormal module $I/I^2$ forces $I$ to be generated by a regular sequence. We prove this implication for every ideal in the linkage class of a complete intersection. When $Q$ is regular, we also establish it under several numerical conditions in terms of the number of generators of $I$ and the type of $R$. For normal domain quotients, we show that finite injective dimension of $I/I^2$ forces $(\mathrm{ht}(I)+1)[ω_R]=0$ in $\operatorname{Cl}(R)$, yielding the desired conclusion when the divisor class group is torsionfree. In the graded setting, we further prove the implication for all monomial ideals in polynomial rings. We also obtain a converse to Kunz's description of the first Koszul homology of an almost complete intersection.

math.AC↗

Biharmonic hypersurfaces in space forms

We prove that every biharmonic hypersurface in a space form of nonpositive sectional curvature is minimal in arbitrary dimension. This settles the hypersurface case of the generalized Chen's conjecture in hyperbolic space and gives a unified treatment of nonpositive space forms. For biharmonic hypersurfaces in the unit sphere, we derive quantitative restrictions on every possible nonconstant-mean-curvature solution. In particular, we obtain pointwise criteria forcing constant mean curvature, a strict scalar-curvature bound in dimensions at least five, and local classification results above the classical CMC gap threshold. These results provide further evidence for the BMO conjecture.

math.DG↗

Optical Signatures of Black Holes Surrounded by a Generalized Cloud of Strings

In this work, we investigate the optical signatures of black holes surrounded by a generalized cloud of strings, described by the Letelier--Alencar spacetime. We first analyze null geodesics, showing that the string parameters can either increase or decrease the photon sphere radius and the critical impact parameter relative to their Schwarzschild values. Since the spacetime is not asymptotically flat, the critical impact parameter does not directly coincide with the physical shadow radius. We use this distinction to derive constraints on the model parameters from the shadow radius bounds of Sgr A* and M87*. In the weak-field regime, we calculate the light deflection angle using both the Gauss--Bonnet method and a perturbative expansion of the null geodesic equations, obtaining consistent results that contain a contribution from the asymptotically conical geometry in addition to the local gravitational deflection. We also examine the Shapiro time delay and identify a distance dependent contribution associated with the non Minkowskian asymptotic background. Finally, we study the optical appearance of geometrically and optically thin accretion disks and construct celestial sphere images through backward ray tracing. The transfer functions and lensing images show that the generalized cloud of strings modifies the null geodesic mapping in a nonuniform manner rather than producing a simple rescaling of the Schwarzschild image. In particular, configurations with a smaller shadow may still produce an outward displacement of the main Einstein ring image, demonstrating that different optical observables probe distinct aspects of photon propagation in this geometry.

gr-qc↗

A relaxation to the post-composition approach to metric-valued Sobolev maps

We establish that the definition of Sobolev mappings from metric measure spaces to metric spaces relying on the post-composition approach of Ambrosio-Reshetnyak admits, under certain transparent assumptions on the spaces, an equivalent reformulation that is substantially weaker than the principal version. In detail, the corresponding definition requires maps of this kind to meet two successive tests in order to be called Sobolev: first a ''qualitative'' one and then a ''quantitative'' one. At the same time, our result allows the relevant conclusion to be drawn based solely on the ''qualitative'' part. Remarkably, the conditions on the spaces sufficient for this to hold fit entirely into the singular flavor of the subject. More explicitly, such a phenomenon emerges if either the source space is of finite Hausdorff dimension and satisfies the Sobolev-to-Luzin-Lipschitz property or the target space is of finite Hausdorff dimension. All this provides a significant generalization of several earlier achievements on the topic.

math.FA↗

Realization of permutation modules as homology of CW-complexes

In this paper, we investigate the realizability problem, which asks whether prescribed group actions on graded modules can be realized by the groups of self-homotopy equivalences of CW-complexes acting on their homology. We provide a partial answer in the case of arbitrary groups and permutation modules concentrated in certain degrees. As a consequence, we realize every group as the group of self-homotopy equivalences of an $R$-local CW-complex with arbitrary prescribed connectivity, where $R$ may be chosen with $ρ(R)$ sufficiently large. In particular, this provides a complete answer to Kahn's realizability problem.

math.AT↗

Note on fractional sum of the divisor function

Let $d(n)$ be the divisor function and denote by $[t]$ the integral part of the real number $t$. We prove that \begin{align*} \sum_{n\leq x^{1/c}}d\left(\left[\frac{x}{n^c}\right]\right)=d_cx^{1/c}+O_{\varepsilon,c}(x^{θ_c+\varepsilon}), \end{align*} where $d_c$ is a suitable constant, $θ_c<2/(3c+2)$ for $0<c<2/11$. This result constitutes an improvement upon that of Y. Feng (2024).

math.NT↗

RoboIRS: Inference-Time Internal Representation Steering for Generalist Robot Policies

Vision-language-action (VLA) and world-action models (WAMs) often degrade under out-of-distribution task variations despite retaining partial task capability. To recover such capability, we propose RoboIRS, an inference-time internal representation steering method that uses successful and failed rollouts to train linear classifiers, select outcome-relevant intervention locations, and derive task-specific steering directions without updating policy parameters. On 15 simulation tasks with a frozen $π0.5$ policy, RoboIRS improves the average success rate from 44.4% to 66.2%, outperforming alternative inference-time intervention baselines while adding little inference time. We further validate RoboIRS on real-robot manipulation using the same $π0.5$ policy and demonstrate its applicability to a world-action model Cosmos Policy, where the average success rate improves from 35.4% to 55.4%. These results show that directly steering internal robot-policy representations can improve the performance of robot policies at inference time. Project website is available at https://rollingoat.github.io/roboirs/.

cs.RO↗

Feasibility of Simultaneous Input-Output Constraints for Tracking in a Class of LTI Systems: Part II

This paper illustrates the properties of a new CBF-governor that enables simultaneous input and output constraint satisfaction for a class of multi-input linear time-invariant systems with state feedback and integral action. This governor is designed so as to modify the reference command while preserving the nominal feedback controller. Necessary and sufficient conditions are derived in a companion paper under which the governor is shown to be feasible at every state of a prescribed operating set and ensures simultaneous input and output constraint satisfaction, forward invariance, and bounded closed-loop solutions. These theoretical results are illustrated in this paper through several numerical examples. In each of these examples, we show how the goals of the CBF-governor are met and corroborate the corresponding necessary and sufficient condition. The computational burden associated with the proposed CBF-governor is also articulated, with its online component comparable to either that of a QP solver or determined using an exact closed-form solution. The offline component is related to the checking of the SIOCF condition which is of $O(N)$.

eess.SY↗

Trends in EHR Satisfaction and Interoperability Among Family Physicians: A Five-Year Analysis of the ABFM Continuous Certification Questionnaire, 2022-2026

Objective: To assess trends in family physicians' (FPs) electronic health record (EHR) satisfaction and their experience of interoperability across organizations. Materials and Methods: Serial cross-sectional analysis of five waves of the American Board of Family Medicine's Continuous Certification Questionnaire (2022-26). Adjusted logistic regression compared odds of being very satisfied across 11 EHRs; Cochran-Armitage tests assessed vendor trends. Ease of using outside information was analyzed separately for same- and different-vendor sources. Results: Across 36,785 FPs, adjusted odds of being very satisfied relative to Epic ranged from 0.91 (95% CI, 0.74-1.11) for Elation Health to 0.17 (95% CI, 0.15-0.20) for Oracle Health/Cerner. Only Epic's users became significantly more satisfied (Benjamini-Hochberg adjusted P < .001), widening the vendor satisfaction gap. Within comparable periods, ease of using different-vendor information did not improve; fewer than 13% rated it very easy. Epic's interoperability advantage reversed by exchange type: non-Epic users had less than half the odds of Epic users of rating same-vendor interoperability very easy (adjusted OR, 0.41; 95% CI, 0.38-0.44) but 1.54 times the odds for different-vendor interoperability (95% CI, 1.35-1.75). Discussion: Rising national exchange volume has not yielded easier use of outside information. Epic's advantage is confined to its own network, so market concentration may shift, rather than solve, the different-vendor problem. Conclusion: The ABFM data are already informing EHR policy but also have utility for the EHR marketplace. While EHR satisfaction has known implications for burden and burnout, poor interoperability is a quality and safety threat that may be compounded by artificial intelligence.

cs.HC↗

Super-Resolution in The Right Latent Space: A Frozen Vision-Foundation Substrate

In an image latent space, the embeddings of high-resolution, natural, and sharp images form a manifold. Degradation of high-resolution images pushes their embeddings off this manifold. Real-world super-resolution (SR) then becomes the task of mapping the degraded embedding back onto this manifold --- not anywhere on the manifold, but to the point that preserves what the input still carries, both its semantics and pixel details. Every published method implements this mapping in a reconstruction-oriented latent space or pixel space. We claim these spaces are the wrong substrates for SR. Low-resolution and degraded images are embedded far from the manifold, making the mapping difficult and expensive. The lack of semantic information in these substrates also makes it difficult to navigate to the faithful point on the manifold, causing severe hallucination when degradation is heavy. Thus, restoring in a suitable latent space is crucial to the SR task. We show that the latent space of 23 fused layers of a frozen DINOv3-L is one such space that makes the SR task easier. Degraded images are embedded near the manifold. Moreover, this substrate contains a hierarchy of information, from pixel record to degradation robust semantics, guiding the model to find the faithful point on the manifold. On this substrate, a 415M decoder is trained under reconstruction and adversarial objectives to map the degraded embeddings back and decode to pixel space in one pass. The resulting model, RAESR, attains the best fidelity--perception trade-off among state-of-the-art adversarial and diffusion-based restorers on RealSR, DRealSR, LSDIR and DIV2K-Val, at 37 ms per 512 by 512 image on a single H20 GPU. Swapping the substrate for a VAE latent under an identical recipe loses on every metric.

cs.CV↗

MemTrace: State-Consistent Memory for Long-Horizon Coding Agents

As coding agents take on long-horizon software evolution tasks spanning multiple files and stages, longer execution trajectories introduce two coupled challenges: (1) accumulated histories strain context budgets, and (2) repository changes can invalidate earlier execution evidence. Existing approaches address these challenges through techniques like larger context windows, compression, retrieval, or repository representations, but often fail to reconstruct a consistent task state after a context refresh or verify whether recalled evidence remains valid. Thus, we introduce MemTrace, a provenance-aware memory system that preserves execution history and aligns its reuse with the evolving task (e.g., iterative cross-file repair) and repository state. MemTrace stores history as immutable Memory Traces anchored to key information (e.g., files, symbols, tests), and organizes their execution order and dependencies in a Memory Trace Graph. When context is constrained, working memory retains only compact Memory Anchors, from which the agent can reconstruct the latest execution state and locate evidence relevant to its next action. Before restoring historical evidence, MemTrace checks its validity against the current repository state and retrieves only what the next action requires. Across three complementary long-horizon coding benchmarks, MemTrace consistently outperforms all fully evaluated baselines under the same backbone and harness, improving DeepSWE pass@1 by 21.2 points, SWE-EVO Resolved Rate by 4.4 points, and SWE-Milestone Score by 17.8 points under Codex CLI.

cs.AI↗

Scaling Verifiable Environments for Long-horizon Work Agents

Work agents operate over digital artifacts to execute professional knowledge-intensive work, requiring training environments that support long-horizon interaction and trustworthy verification. However, hand-crafted environments incur prohibitive engineering overhead that prevents environment scaling, whereas synthesis methods sacrifice workspace complexity, realism, or grounded verifiability. To bridge this gap, we introduce WorkForge, a scalable synthesis framework for constructing verifiable work-agent environments from real-world resources. Starting from expert workflows, WorkForge first identifies the resources, decisions, and deliverables required by each workflow. It then retrieves relevant real-world files and organizes them into a workspace. WorkForge inspects the workspace to extract concrete, checkable facts about its content. These factual anchors fix which task types the workspace can support and how their outcomes can be verified. Therefore, WorkForge derives each task's instructions, solution plan, and complementary programmatic and semantic verifiers directly from these factual anchors, keeping verification traceable to observable workspace evidence. Furthermore, we construct 16.7K verifiable environments across 40 professional domains, with workspaces collectively covering 60 file types. Post-training Qwen3.5-35B-A3B-Base improves GDPVal from 45.5 to 73.6 and APEX Score from 5.0 to 21.3, while enabling Qwen3.5-27B to achieve highly competitive performance and outperform strong competitors. Our analyses confirm the efficacy of the proposed method and reveal consistent scaling behaviors across both data volume and interaction horizons.

cs.CL↗

AlphaPADI: Formulaic Alpha Discovery via Pool-Aware Hierarchical Discrete Diffusion

Formulaic alpha discovery seeks symbolic expressions that predict cross-sectional asset returns. In deployment, multiple formulas are combined into an alpha pool, where each formula is valued through the complementary information it contributes to joint predictive performance. While Reinforcement Learning and Generative Flow Networks have emerged as promising paradigms for generating formulaic alphas, existing frameworks face three related challenges. First, generating formulas individually leaves pool context and inter-formula complementarity outside the generative state. Second, formula-wise generation lacks a unified mechanism for preserving and revising structures at different levels. Third, pool-level rewards jointly reflect predictive performance and redundancy but cannot be differentiated directly through symbolic evaluation to train the generator. To overcome these challenges, we introduce AlphaPADI (Formulaic Alpha Discovery via Pool-Aware Hierarchical Discrete Diffusion), a novel framework built around three components: (1) grammar-constrained buffer initialization that constructs syntactically valid pool candidates, (2) pool-aware hierarchical diffusion that reconstructs complete pools at multiple structural scales under the current pool context, and (3) reward-guided pool refinement that evaluates joint predictive performance and inner diversity, updates the elite buffer, and trains the reverse model through reconstruction and preference learning. Empirical results on the Chinese and U.S. stock markets demonstrate that AlphaPADI outperforms the evaluated baselines in both predictive and portfolio performance, thereby validating pool-aware generation as an effective framework for automated alpha discovery.

cs.CE↗

FreeLoc: Online Floorplan Localization via Diffusion-Aided Pose Refinement

Floorplans provide compact and widely available geometric maps for indoor localization, but existing high-performing floorplan-based methods still convert them into dense scene-specific offline databases, tying accuracy, storage, and runtime to the sampling resolution of the discretized pose space. We present FreeLoc, an online RGB-based floorplan localization framework that treats the floorplan as a directly queryable geometric map. FreeLoc introduces an efficient online geometric querying and diffusion-aided refinement scheme, which retrieves plausible pose anchors through on-the-fly floorplan ray querying and refines them into accurate continuous pose estimates. For sequential localization, FreeLoc develops an online likelihood construction strategy that bridges single-frame localization and probabilistic temporal fusion by constructing likelihoods from coarse-sampled candidates and refined pose hypotheses, enabling histogram-filter-based temporal fusion without offline databases. Experiments demonstrate real-time online inference and state-of-the-art performance in both single-frame and sequential localization, while real-world results validate practical deployability in indoor robotic localization scenarios.

cs.RO↗

Measuring Learned Monotone Temporal Aggregation at Matched Admissibility

Risk regulation imposes directional constraints on scores; we adopt their strict per-input form -- the score monotone non-decreasing in every exposure input -- as a normative commitment. Deployed pipelines -- monotone hand-crafted aggregates feeding sign-constrained gradient boosting -- already satisfy it by composition, so constrained-versus-unconstrained comparisons price a guarantee the incumbent has for free. We instead hold admissibility fixed on both sides and measure what learning the aggregation is worth. Our instrument is a recurrent network whose state is classical risk statistics (an exponentially weighted moving average and a high-water mark with learned transforms), monotone by construction in every input and per MC-dropout sample. The central finding, by functional regression, is a subsumption boundary: a learned monotone channel reproduces the geometrically weighted separable family of hand-crafted statistics, one channel per member, to Spearman $ρ\ge 0.996$, approximates window statistics with measurable ceilings, and fails at consecutivity ($ρ= 0.924$) and time localization (0.628), both structural, and at the exposure floor (0.829), a learnability boundary. One explicit admissible basis repairs each failure (rank correlation 1.000). In or near the separable family, learned and engineered aggregation are substitutes, and the learned channel is never statistically behind at full sample size and specified capacity. Its advantages are incumbent-specific: a committed grid pays up to 0.019 AUC in decay regions it leaves uncovered (the learned channel stays within 0.004 of the strongest engineered consumer at every swept point); the highest-dimensional comparator degrades fastest with scarce data; and beyond the training support, grid-fed tree-ensemble scores go flat while a strictly increasing head keeps ranking. No single incumbent is dominated on all three axes.

cs.LG↗

Understanding the Hierarchical Structure and Functional Landscape of the Model Context Protocol Ecosystem

AI agents increasingly rely on tools exposed through the Model Context Protocol (MCP) to complete user tasks. Hundreds of thousands of MCP servers are listed across marketplaces, yet they are organized only by coarse, marketplace-specific server categories. This makes it difficult for agents and users to identify tools for a given operation, find functional alternatives, and assess how those alternatives differ. We present MCPacific, the largest tool-level, cross-marketplace map of the MCP ecosystem. MCPacific collects 368,754 MCP server listings corresponding to 124,267 unique servers across 17 marketplaces, statically extracts 1,328,233 tool specifications from these servers in seven languages, and organizes them into a hierarchical functional taxonomy of 58,915 capabilities. We construct the taxonomy through an iterative LLM-driven design-test-refine process and map the full corpus to it using calibrated embedding routing. Our study reveals that MCP extends well beyond developer tooling, with 85% of tools serving other domains. Functional alternatives are widespread but unevenly distributed: 98.5% of tools have at least one alternative, yet nearly a quarter of capabilities are supported by only one tool. Functionally comparable tools also differ in security alerts, code complexity, and project maintenance, with complexity differing by more than 2.5x in 41% of comparable tool pairs. Finally, presenting candidate tools through the taxonomy rather than a flat list improves task completion rate across all four evaluated models, with gains of up to 12 percentage points in Pass@0.75 for crowded candidate sets.

cs.SE↗