Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 433 records · Page 24Linked to original sources

Beyond Plausibility: Verifiable Fine-Grained Image Editing on Structured Assets

Fine-grained image editing requires more than producing a visually plausible result: an editor must execute the requested attribute change precisely while leaving everything else intact. However, existing benchmarks leave a critical gap between realism and verifiability: benchmarks built on realistic images typically rely on human or vision--language model judgments, while deterministic evaluation has largely focused on synthetic shape canvases, with application-oriented extensions primarily limited to charts. This makes it difficult to determine precisely how much of a requested edit was executed, where unintended changes occurred, and whether small differences between models reflect genuine editing capability or evaluator uncertainty. To bridge this gap, we present VeriEdit-Bench, a benchmark for fine-grained, instruction-faithful image editing across realistic structured assets with deterministic, four-axis evaluation. Its 1,740 cases are compiled from the source code of 153 Scalable Vector Graphics (SVG) graphics, charts, web interfaces, and presentation slides. Controlled source-code edits preserve the original visual context while yielding exact target images, pixel-level edit masks, and explicit edit specifications, enabling reproducible scoring along four axes: edit fidelity, preservation, localization, and magnitude. Evaluating eleven editors, we find that even the strongest model remains far from full credit; rankings for the same recoloring operation reverse between charts and SVG graphics; and outputs with similar pixel-accuracy profiles can still differ substantially in localization and change magnitude. This decomposition yields graded, verifiable feedback and exposes model-specific capability and failure profiles that holistic scores or evaluator-dependent judgments may obscure.

cs.CV↗

Mixed configuration spaces, fixed points and Loop braid groups

We introduce the notion of mixed configuration spaces of the $3$-ball. That is, we consider both distinct circles and distinct points in the interior of the $3$-ball and we prove that the projection onto the configuration space of circles of the $3$-ball is a locally trivial fibration. The key idea for defining the mixed configuration spaces is to study fixed points of orientation-preserving homeomorphism of the $3$-ball that leave invariant a trivial link of $n$ components in the interior of the $3$-ball. In particular, we prove that two fixed points are Nielsen equivalent if and only if the associated loop braids are conjugate by an element of a distinguished free subgroup of rank $n$. This result stands as a $3$-dimensional counterpart of the $2$-dimensional result where fixed points of homeomorphisms of the punctured disc are characterized in terms of braid elements of the classical Artin braid group.

math.GT↗

Conditioning and Singular Value Distribution of Jacobi Weighted Histopolation Matrices

In this paper, we first study the singular value distribution of the unscaled matrices arising from Jacobi weighted histopolation. For the Jacobi weight with parameters $(α,β)$, and using the Jacobi basis with parameters $(α+1,β+1)$, we previously proved that, for $α,β>1/2$ and quasi-uniform meshes, the growth of these matrices in the sense of singular value distribution is slower than any positive power of the logarithm of the matrix size. However, the numerical behavior of the unscaled matrices suggested a nontrivial singular value symbol, which we identify explicitly in this paper. Starting from the standard asymptotic formula for Jacobi polynomials, we derive an asymptotic representation of the weighted cell averages defining the entries of the histopolation matrices. This representation allows us to analyze the associated Gram sequence by means of GLT theory. We identify its GLT symbol and, consequently, we obtain the singular value distribution of the histopolation matrices for any $α,β>-1$. We also characterize the set $\mathcal Z$ where the resulting symbol vanishes and compute its measure, showing that this set depends only on the mesh, while the Jacobi parameters affect the values of the symbol on its support. When this measure is positive, the singular value distribution implies the discrete asymptotic ill-posedness of the problem, while a comparison with the Toeplitz setting suggests exponential growth of the spectral condition number. When the measure of $\mathcal Z$ is zero, the problem is well-posed and the observed conditioning is algebraic in the matrix size. Finally, numerical experiments supporting the theoretical findings are reported and discussed, while conclusions are given in the final section of the current work together with a few open problems.

math.NA↗

A New insight into the Ambrosio--Reshetnyak approach to Sobolev maps

Given $p \in (1,\infty)$, let $\operatorname{X}=(\operatorname{X},ρ,μ)$ be a metric measure space such that the measure $μ$ is uniformly locally doubling and $\operatorname{X}$ supports a weak local $(1,p)$-Poincaré inequality. Let $(\operatorname{Y},\operatorname{d},\underline{y})$ be a complete pointed metric space. We prove that the equivalence class of a Borel map $ u:\operatorname{X} \to \operatorname{Y}$ (modulo coincidence $μ$-a.e.) belongs to the Ambrosio--Reshetnyak--Sobolev class $W_{p}^{1}(\operatorname{X},\operatorname{Y})$ if and only if $h \circ u$ belongs to the Sobolev space $W_{p}^{1}(\operatorname{X})$ for every $1$-Lipschitz function $h:\operatorname{Y} \to \mathbb{R}$.

math.FA↗

Largest Rashomon sets of decision trees for robust contextual optimization

Many decision trees fit the same data almost equally well, yet they can route a query point to different leaves and induce different local empirical distributions. We study decisions that meet prescribed cost, shortage or risk targets despite this predictive multiplicity. We propose the joint Rashomon and robustness optimization framework for optimal decision trees. It jointly selects an operational decision and the largest Rashomon set of trees, so that the targets hold under the local empirical distribution that every tree in this set induces at the query point. We specialize the framework to the regression setting, and we show that a tree affects the decision only through the training observations sharing the query leaf, which we call its query neighborhood. As a result, the robust problem involves only finitely many distinct constraints, which can be examined in order of increasing estimation loss. We develop a constraint generation algorithm that combines query-path pricing with dynamic programming to identify violating neighborhoods without enumerating trees. On synthetic newsvendor instances, the algorithm typically needs few neighborhoods and runs substantially faster than full neighborhood enumeration. On restaurant demand data, the robust orders increase the mean tolerated excess estimation loss by 17.6% and reduce the empirical conditional value-at-risk of the worst 10% of realized costs by 8.3% relative to the sample average approximation orders of the optimal tree, while the mean cost difference is not statistically significant. An interpretability analysis further shows how the retained neighborhoods explain the decision and its robustness limit.

math.OC↗

Pion gravitational form factors in an algebraic rainbow-ladder model

Within rainbow-ladder truncation, we identify three contributions to the pion's gravitational form factors---graviton coupling to the quark, to the gluon dressing the quark, and to the gluon binding the pion---and separate them into quark and gluon parts. Using an algebraic quark model with a dynamical basis, we obtain $A(0)=1$, $D(0)=-1$ in the chiral limit, and momentum fractions $\langle x\rangle_q=0.896$, $\langle x\rangle_g=0.104$. The D-term obeys the same partition at the model scale.

hep-ph↗

TrustMed-RL: Long-Horizon Reinforcement Learning for Evidence-Grounded Clinical Diagnosis

Medical language models can produce correct diagnoses despite incomplete investigations and unsupported reasoning. To support long-horizon, evidence-grounded diagnosis, we introduce \textbf{TrustMed-RL}. Built from PubMed rare-disease cases and over 24,000 manually annotated image panels, it integrates interviews, examinations, testing, specialist consultation, and literature search through state-dependent actions. Our 8B vision--language policy, trained with clinically adapted GiGPO and coverage-adjusted diagnostic rewards, achieves 37.1\% diagnostic accuracy on 2,500 evaluation cases, outperforming all evaluated open-weight baselines and improving over supervised fine-tuning by 12.4 percentage points.. When success additionally requires acquiring at least 50\% of supporting test evidence, TrustMed-RL achieves 32.5\%, exceeding GPT-4o by 6.8 percentage points. Furthermore, it surpasses all evaluated baselines on MTMedDialog and multiple larger 27--32B models on AgentClinic. In physician review of 200 diagnostically accepted test-set trajectories, 83.0\% receive evidential-grounding scores of 4--5 out of 5. Physicians' assessments suggest that these diagnostic trajectories are trustworthy and aligned with human diagnostic reasoning.

cs.AI↗

Kilovolt-class Vertical (011) \b{eta}-Ga2O3 Schottky Diodes for High Voltage and High Temperature Applications

We report kilovolt-class vertical (011) \b{eta}-Ga2O3 Schottky diodes utilizing high-permittivity dielectric TiO2/Al2O3 field-plate for high voltage and high temperature applications. A systematic study was performed with Pt/(011) \b{eta}-Ga2O3 Schottky barrier diodes (SBDs) with varied diameters from 100 μm to 300 μm, for both with and without field-plate, that revealed excellent consistency of reverse blocking performance regardless of diode area. The field-plate SBDs achieved superior breakdown voltages (3.90 - 4.16 kV) compared to the diodes without field-plate (3.06 - 3.16 kV) owing to effective edge termination. Furthermore, we explored high-temperature performance of the field-plate (011) \b{eta}-Ga2O3 diodes that revealed excellent forward conduction properties and high rectification ratio (~10^9) throughout the temperature range (25 - 150 C). At reverse bias, the field-plate diodes also exhibited kV-range breakdown voltage with no evident increase in leakage current up to the explored elevated temperature of 150 C. The low area-dependence of breakdown voltage, excellent forward transport properties, and minimal reverse leakage at both high-voltage and elevated temperature demonstrate the exciting potential of (011) \b{eta}-Ga2O3 SBDs in high-temperature power switches. Thus, our work demonstrates the strategy of advancing the performance of vertical \b{eta}-Ga2O3 SBDs by utilizing the advantageous (011) \b{eta}-Ga2O3 epilayers with low background doping, reduced killer dislocation effects, and effective field management for high-voltage and high-temperature applications.

cond-mat.mtrl-sci↗

Arbitrarily High-Order Structure-Preserving Parametric Finite Element Methods for Geometric PDEs via Local Pullback Flow Maps

We develop arbitrarily high-order structure-preserving parametric finite element methods for geometric PDEs including surface diffusion and volume-preserving mean curvature flow based on local pullback flow maps. The central idea is to reformulate the geometric PDEs on an arbitrary intermediate hypersurface and represent the subsequent evolution through a local pullback flow map. Using admissible local pullbacks, we derive two formulations from the corresponding pullback identities for the Laplace-Beltrami operator, referred to as the conformal and equidistribution formulations. We discretize both formulations using arbitrary-degree isoparametric finite elements and Runge-Kutta collocation methods with positive weights. For the conformal formulation, weighted averages of the normal-Jacobian vector ensure exact preservation of the enclosed area or volume, while algebraic stability of the Runge-Kutta method additionally guarantees dissipation of the perimeter or surface area. For the equidistribution formulation, a path-averaged normal-Jacobian vector and a geometric discrete gradient guarantee both exact preservation of the enclosed area or volume and dissipation of the perimeter or surface area, without requiring algebraic stability. Numerical experiments confirm the expected high-order accuracy and structure-preserving properties.

math.NA↗

Gaussian Flow Dynamics: Simulation-Free Neural SDE Learning Beyond One-Time Marginals

Simulation-free training of latent Stochastic Differential Equations (SDEs) relies on a variational posterior process whose one-time marginals are tractable, typically Gaussian. Such marginals, however, do not determine the underlying dynamics: many processes share the same marginals while differing in their temporal structure, and existing parameterizations fix this structure implicitly, which restricts the posterior family and biases the learned model. We introduce Gaussian flow dynamics, which construct stochastic processes directly from smoothly evolving Gaussian marginals while making the marginal-preserving, or gauge, degrees of freedom explicit and parameterizable. The construction admits state-dependent diffusion coefficients and recovers every linear SDE with additive noise and a non-degenerate Gaussian initial distribution. Building on it, we propose Gauge Matching, a simulation-free method for latent SDE learning that combines Gaussian flow dynamics with the SDE Matching objective. Gauge Matching costs at most quadratically in the latent dimension per step, like SDE Matching, but learns the temporal structure of the posterior beyond its one-time marginals. It comes within a nat of Helmholtz-SDE, which computes the gauge from the prior Jacobian at cubic cost, on the linear benchmark where the exact posterior is known, matches it on nonlinear systems, and applies where Helmholtz-SDE does not, to state-dependent noise.

stat.ML↗

AgenticTactileVLA: Contact-Guided Execution-Time Supervision for Generalizable Dexterous Manipulation without VLA Retraining

Vision-language-action policies may predict a transferable manipulation strategy yet fail to realize it reliably on the encountered object: objects compatible with the same grasp differ in geometry and compliance, and visual feedback degrades under closure occlusion. AgenticTactileVLA is presented as an execution-time supervisor that shifts part of object-specific adaptation from prediction to physical interaction. A fixed VLA provides the approach and hand targets; the supervisor decides whether to remain transparent, refine finger flexion, retain or release the corrected configuration, return control to the VLA for retry, or select a compliant hand-control regime. It uses finger-position and motor-effort feedback as proprioceptive contact evidence and requires neither tactile sensors nor VLA retraining. On a Unitree G1 with a BrainCo Revo2 hand, a randomized matched-block evaluation on five objects held out from VLA training yields 61.3% completion for the base VLA, 72.0% for unconditional close-to-stall control, and 84.0% for the supervisor under a shared budget; the gain is positive on every object and persists under moderate pose perturbations. Ablations show the gain is not explained by extended closure alone, and that selective triggering reduces correction episodes by 65.7% with no detected change in completion. A retention audit shows acceptance predicts retention in 88.9% of held-out cases, while compliant objects expose conservative false rejection. A thin-walled-cup study demonstrates contextual routing to compliant control, matching an always-compliant reference. These results suggest that contact-guided execution-time adaptation can improve the object-level generalization of a fixed VLA to held-out objects by adapting physical realization without object-specific retraining.

cs.RO↗

Detecting Defects that Matter: An Application-Driven Benchmark for Anomaly Detection in Manufacturing and Retail Logistics (VAND 4.0 Challenge)

Existing Anomaly Detection benchmarks are saturated and often unrealistic. As part of the VAND 4.0 Challenge, we introduce a hidden-test, application-driven benchmark across two deployment-critical domains: industrial manufacturing and retail logistics. In the Industrial Track (MVTec AD 2), the results reveal that unsupervised anomaly segmentation remains challenging: the best regular-setting method achieves only ~57\% pixel-level $SegF_1$, indicating substantial room for improvement. Zero-shot approaches trail by ~15 $SegF_1$ points, confirming that task-specific training on normal data remains essential for precise defect localization. Robustness to distribution shifts remains a key open challenge and DINOv3-backbones clearly dominate this track. In the Retail Track (Kaputt 2), the results reveal that (1) supervised defect detection is approaching saturation for common defect types; (2) the best off-the-shelf VLM approach trails specialized models by ~28 AP, confirming that currently VLMs cannot replace fine-tuned detectors, (3) reference images did not prove helpful for top-performing approaches. Performance collapses on rare defects (spillage ~53 AP, missing units ~27 AP), where the supervised ceiling is bounded by data availability. To drive future progress in this domain, we provide a new low-prevalence retail AD dataset (Kaputt-Rare). Across both tracks, computational efficiency is assessed as a first-class metric combining performance, throughput, memory, and power consumption. We introduce a novel metric for measuring efficiency and reveal that that top-performing methods rely on heavy architectures while efficiency is largely neglected. Overall, we conclude that the community needs (a) more efficiency-aware method development, and (b) true anomaly detection approaches for rare defects and shifting conditions. https://sites.google.com/view/vand4-cvpr2026/challenge

cs.CV↗

Structural Conditions for Distributed Quantum Advantage

Circuit cutting runs a quantum computation on processors smaller than the circuit, at the price of classical knitting whose cost grows exponentially with the number of cut gates. Keeping this cost affordable is necessary but not sufficient for an advantage, since knitting classically easy subcircuits is itself classically easy. We formulate three requirements for distributed quantum advantage under a stated budget: affordable knitting, classical difficulty that survives the cut, and, for variational circuits, resolvable gradients. We prove that under stated assumptions knitting stays affordable when subcircuits grow behind a bounded interface and becomes exponentially expensive when fixed-width subcircuits multiply. We turn the requirements into screening criteria and apply them to eighteen circuit families. We identify finite local-depth circuits with bounded interfaces as a setting in which the requirements can coexist, with hardness established only in the worst case. As a classically verifiable proof of principle, we knit one gate joining two field-perturbed toric-code patches on an IBM Nighthawk processor for up to 142 spins, and confirm that the reconstruction can at least resolve the contact correlation beyond independent execution for up to 98 qubits.

quant-ph↗

A Self-Generated Local Structural Prior for Resolving Target Merging in Electrical Impedance Tomography (EIT)

Electrical impedance tomography (EIT) is an ill-posed inverse imaging technique with limited and spatially nonuniform resolution. Conductivity targets located in regions with degraded spatial sensitivity may be reconstructed as a single connected region, resulting in the loss of spatial separability between neighboring structures. In this work, we propose a self-generated local structural prior for suppressing target merging in EIT reconstruction. First, an initial conductivity image is obtained from the measured boundary data using an existing EIT reconstruction method. This reconstruction is subsequently used as a reference image to construct a Gaussian similarity matrix according to the intensity similarity between image pixels. The adaptively chosen eigenvector of the covariance matrix is then employed to extract the principal structural information of the reference image. A localized structural mask is constructed and incorporated into a subsequent EIT reconstruction to preserve the spatial separability of neighboring targets. Numerical and water tank experimental results demonstrate that the proposed method effectively separates neighboring conductivity targets that are merged by conventional EIT reconstruction methods.

math-ph↗

Projection theorems for macroscopic intermediate dimensions

We introduce lower and upper macroscopic intermediate dimensions of unbounded subsets of Euclidean space by combining covering contributions from all dyadic annuli beyond a given scale. For every unbounded Borel set $E\subset\R^d$ and $m\in \{1,\ldots,d-1\}$, we prove that the macroscopic intermediate dimensions of $P_VE$ equal the corresponding intrinsic capacity profiles for almost every $m$-dimensional linear subspaces $V\subset \R^d$. One full-measure set works for both dimensions and every parameter $θ\in[0,1]$. The proof combines a covering--capacity comparison with an exponent gap, fractional-moment estimates for projected coverings, and capacity-weighted retention of source shells. The dimensions recover macroscopic Hausdorff dimension at parameter zero, are invariant under quasi-isometries, and are continuous at positive parameters. The upper macroscopic intermediate dimension agrees with the upper discrete intermediate dimension introduced in \cite{LXZ}, whereas the lower one can differ. We compute the projected dimensions of bounded-base discrete digit Moran sets, including the self-similar case obtained when the bases and digit sets are constant, and give applications to Lipschitz graphs and arithmetic sumsets.

math.CA↗

JASPER: Special Session on Joint Reliability And Security Assessment of SPlit Computing for Edge Robustness

Split Computing (SC) enables efficient deployment of Deep Neural Networks (DNNs) by partitioning inference between edge devices and cloud servers. However, intermediate feature representations are simultaneously exposed to hardware faults and adversarial attacks, which are traditionally evaluated independently. This paper presents a unified framework for the joint assessment of reliability and security in Split Computing. First, reliability is characterized through neuron-level fault injection using the Mean Relative Accuracy Degradation (MRAD) while security through feature-map-aware adversarial attacks simulations using the Attack Success Rate (ASR). Based on these complementary analyses, the Joint Vulnerability Score (JVS) is introduced, along with a confidence-aware extension that jointly captures prediction errors and confidence degradation. The framework is evaluated on ten Split Computing configurations based on ResNet-50 trained on ILSVRC-2012. Experimental results show substantial differences across compression strategies, with MRAD ranging from 44.3% to 61.2% under fault injection, while adversarial attacks achieve up to 98.8% ASR. Furthermore, the proposed joint metrics reveal vulnerability trends that remain hidden when reliability and security are analyzed independently, providing a more comprehensive methodology for designing dependable Split Computing systems.

cs.CR↗

Simultaneous measurement of the Higgs boson decay rates to WW and ZZ in fully hadronic final states at FCC-ee

The precise determination of the Higgs boson couplings to electroweak gauge bosons is a key objective of the FCC-ee physics program. In particular, measurements of the Higgs boson couplings to W and Z bosons play a central role in the model-independent determination of its total decay width. This paper presents a study of Higgs boson decays to H -> WW and H -> ZZ in the Higgsstrahlung production process. The analysis focuses on events in which the associated Z boson and the vector bosons from the Higgs decay all decay hadronically, resulting in six-jet final states. The large H -> WW branching fraction leads to high signal statistics in the fully hadronic channel, while for H -> ZZ the much smaller Higgs branching fraction is partly compensated by the large hadronic branching fractions of the Z bosons. Due to the substantial kinematic overlap between the WW and ZZ decay modes, their signal strengths are extracted simultaneously. Jet-clustering and dedicated jet-pairing techniques are employed to optimize the reconstruction of the hadronically decaying gauge bosons and enhance the separation of the two signal components. Assuming a center-of-mass energy of 240 GeV and an integrated luminosity of 10.8 ab^-1, expected relative uncertainties of 1.48% and 8.20% are obtained on the production cross section times branching fraction, sigma(ZH) x BR, for the H -> WW and H -> ZZ decay modes, respectively.

hep-ex↗

Effective Dimensions in Grothendieck--Hölder Inequalities

What becomes of Grothendieck's inequality when its Hilbert-space vectors are replaced by vectors in $\ell_p^r$ and $\ell_{p'}^r$? For an $m\times n$ matrix, we show that the universal dimensional loss is governed not by $r$ alone, but by the effective dimension \[ d=\min\{r,m,n\}. \] More precisely, the optimal constant is bounded above and below, up to dimension-free factors, by \[ d^{|1/p-1/2|}. \] The upper bound rests on a pointwise Hilbertian stabilization: for every matrix $A$, \[ G_2^{(r)}(A)=G_2^{(d)}(A). \] Combining this stabilization with Lewis's Euclidean distortion estimate gives the effective-dimensional Grothendieck--Hölder bound. Beyond the universal power law, we retain the finer rectangular geometry. At $p=\infty$ we prove the exact compression formula \[ Γ_\infty^\K(r;m,n) = ρ\!\left( \ell_1^m(\K), \ell_1^{\min\{r,n\}}(\K) \right), \] and derive from it endpoint-sensitive lower bounds throughout the full Hölder scale. Over the complex field, interpolation between the Hilbertian problem and the exact rectangular endpoints yields corresponding upper bounds. Thus the effective dimension determines the universal exponent, whereas the tensor-norm endpoint retains additional aspect-ratio information. We determine the boundedness threshold when $p$ approaches $2$: it is governed by \[ |1/p-1/2|\log d. \] For complex two-row matrices, the endpoint tensor ratios are explicit, and the general rectangular estimates give quantitative two-sided bounds for every $1\le p\le\infty$.

math.FA↗