Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

4DMulti: automated multicomponent identification at complex material interfaces

Mapping crystalline phases at heterogeneous interfaces is essential for understanding material performance and degradation. However, structural heterogeneity, phase overlap, and local disorder complicate diffraction interpretation, while growing data volumes make manual analysis increasingly impractical. We introduce 4DMulti, a physics-guided learning framework for automated multicomponent identification from large-scale four-dimensional scanning transmission electron microscopy (4D-STEM) data. The supporting diffraction data resource comprises over 6 million high-quality experimental patterns and labeled patterns generated by Sim2real. A retrieval-conditioned latent diffusion transformer (Sim2real) translates simulated patterns into experimental-style examples under constraints designed to preserve Bragg geometry, while a rotation-invariant coordinate convolutional network identifies phases across in-plane rotations. 4DMulti achieves 98.82% classification accuracy on a five-phase experimental nanoparticle benchmark, with ablation studies supporting the complementary benefits of domain adaptation and rotation-invariant classification. We define diffraction-inferred structural complexity (DISC), a normalized predictive entropy score that quantifies phase-assignment ambiguity within a specified candidate phase library. We apply 4DMulti to generate structural maps of superconducting heterostructures, corroded alloy surfaces, and degraded solid-state battery interfaces down to single-nanometer spatial resolution. 4DMulti connects simulation-derived crystallographic knowledge to automated experimental interpretation, establishing a foundation for scalable analysis of complex interfaces and data-driven discovery of interfacial design principles.

cond-mat.mtrl-sci↗

A proximal-linearized NEP method for composite optimization with conic and manifold constraints

This paper studies difference-of-convex (DC) composite optimization problems with conic and manifold constraints. By penalizing the conic constraint with a distance-based penalty, we propose an inexact proximal-linearized nonsmooth exact penalty (iPLNEP) algorithm. The proposed method successively finds approximate minimizers of proximal-linearized subproblems over the tangent spaces of the manifold based on computable inexactness criteria, while adaptively updating the proximal and penalty parameters. Under a boundedness assumption on the iterate and penalty parameter sequences, iPLNEP is shown to achieve an $O(ε^{-2})$ worst-case iteration complexity bound for finding an $ε$-stationary point. Using accelerated semi-proximal ADMM to solve the subproblems, we establish an overall oracle complexity bound of $O(ε^{-4})$. Under an additional uniform error-bound condition, using semi-proximal ADMM as the subproblem solver yields an improved overall oracle complexity bound of $O(ε^{-2}\logε^{-1})$. To the best of our knowledge, this provides the first proximal-linearized NEP framework with provable oracle complexity guarantees for DC composite optimization with conic and manifold constraints. If, in addition, the associated potential function satisfies the KL property, the whole sequence of iterates converges to a stationary point. Extensive numerical experiments on composite optimization problems with orthogonal-manifold, nonnegative cone, and second-order cone constraints demonstrate the effectiveness of the proposed method.

math.OC↗

Dynamical Readout of Measurement Statistics and Emergent Entanglement-Like States in Classical Networks

Can a classical network not only encode a quantum-like state, but also read out its measurement statistics through its own collective dynamics? Here we introduce a network-native readout scheme for a structured classical network whose community modes form an effective two-qubit state space. Network connectivity selects a dominant collective mode that encodes the state, while a fixed set of ten elementary connectivity perturbations probes its collective response. The resulting shifts of the dominant growth rate form a complete basis of expectation values for the real two-qubit sector, from which joint-outcome probabilities associated with arbitrary real projectors are reconstructed by linear combination. Once these ten spectral responses are measured, the same data generate correlations over a continuous family of measurement settings. As a benchmark, for an encoded Bell state four selected correlations give $|S|=\chsh>2$, whereas a separable product-state reference remains within the Clauser--Horne--Shimony--Holt (CHSH) bound $|S| \leq 2 $. These results establish an informationally complete dynamical readout of effective two-qubit states directly from the spectral response of a classical network.

quant-ph↗

PaxosLease Revisited: A Checked Model of Diskless Distributed Leases

PaxosLease is a protocol by which a quorum of acceptors grants time-bounded exclusive ownership with no durable acceptor lease state and no disk write on the lease acquisition path. This paper gives a precise, machine-checked statement of the protocol and of its standard use, electing a Multi-Paxos leader. Formalizing and model checking the original protocol changes three rules of its acceptor: two are required for safety, the third allows shorter restart quarantines. The protocol is formalized in TLA+ and checked by TLC, its timing arithmetic is proved in TLAPS, and an executable Python model demonstrates the distributed algorithm for human readers.

cs.DC↗

Protected tails and polynomial-time enumeration of permutations avoiding a direct sum of an increasing pattern and 231

We give an algorithm counting the permutations that avoid a fixed pattern of the following form: the direct sum of an increasing pattern and 231. The first members of the family are 1342 and 12453. For each member the algorithm uses polynomially many operations and stored integers, with degrees that grow linearly in the length of the pattern. It comes from a recurrence that reads a permutation from left to right and records the constraints that the letters read so far impose on those still unread. This recurrence has exponentially many states, but part of each state is protected: later steps carry it along unchanged and do not depend on it, and factoring the protected part out leaves a dynamic program of polynomial size. For 12453 a translation symmetry sharpens the bounds to degree seven for the operations and degree four for the storage. We also compute the number of 12453-avoiding permutations of every length up to 150. The previously published series, due to Biers-Ariel (2019), reached length 38. We also give a sampler of uniformly random avoiders. A floating-point implementation of it, proved to be within total variation distance $3.5\cdot10^{-5}$ of uniform for ideal random bits, draws the one million 12453-avoiding permutations of length 300 shown in a heatmap. The literal and kernel recurrences for 1342 and 12453 are verified in the Lean 4 proof assistant.

math.CO↗

Anchored Sequential Deliberation

Sequential deliberation is a mechanism for collective decision making: at each round, a uniformly randomly selected pair is asked to revise a collective outcome, which then becomes the input for the next round. Existing theory by Fain et al.~\cite{fain2017sequential} treats the current outcome solely as the disagreement alternative in bargaining. Yet the existing outcome might also carry social influence and anchor participants' positions toward the status quo. In this paper, we introduce \emph{anchored sequential deliberation}. At each round, two participants with bliss points $U$ and $V$ shift their positions toward the previous outcome $O_{t-1}$ with anchoring strength $0\leq λ<1$. They then bargain using $O_{t-1}$ as the disagreement alternative. For the sake of analysis, we assume that the decision space is one-dimensional, the anchoring effect is linear, and participants use Nash bargaining. The update simplifies to $O_t=(1-λ)Med\{U,V,O_{t-1}\}+λO_{t-1}$. Our analysis reveals a trade-off. Through a coupling of two outcomes, we find that the process contracts in $1$-Wasserstein distance with a factor of at most $\frac{1+λ}{2}$, which implies that stronger anchoring slows mixing. On the other hand, stationary distortion weakly decreases with $λ$, although the worst-case distortion remains $\frac{1+\sqrt{2}}{2}$ for every feasible $λ$. We also identify a unique \emph{deliberative fixed point}, at which the expected movement is zero, and prove that the stationary distribution concentrates around it as $λ$ increases. For symmetric populations, we further provide a tighter bound on the stationary variance around this fixed point. For the uniform distribution, we establish upper and lower bounds on stationary distortion, both of which approach $1$ as $λ$ increases.

cs.GT↗

De-GAN - Dynamic Parameter Tuned GAN for 3D Medical Image Segmentation: A Step Towards Generalisation

Brain tumor segmentation remains difficult because enhancing tumor (ET) has low contrast and overlaps surrounding tissue, while scanner and site variation causes domain shift. We propose DE-GAN, a contrast-enhancing conditional GAN that combines input-adaptive dynamic convolutions, style-aware feature mixing, and coordinate encoding to synthesize slice-adaptive FLAIR images. A label-guided, class-conditional target separates tumor-core (TC) and ET intensities while preserving anatomy. The generated FLAIR is concatenated with the original MR modalities and used to train a 3D U-Net. Across BraTS 2015, 2018, and 2019, DE-GAN improves segmentation over the baseline and static EnhGAN replacement on most reported TC/ET metrics, with the largest gains from retaining both original and enhanced FLAIR. Code and pretrained models are available at https://github.com/zkhansuri-ui/DE-GAN.

cs.CV↗

A Data-free Universal Prior over Syntactic Structures

The probabilities of syntactic structures in human languages are assumed to emerge fully from language-specific experience. Here, I show that a universal prior over syntactic structures emerges from a model of human language production, in which words are progressively integrated into syntactic structure. Without fitting any parameters to specific language data, the resulting prior assigns higher probabilities to attested than to random dependency trees in all 138 typologically diverse languages examined. These prior probabilities correlate positively with those estimated from corpora in 33 of 34 languages. The results indicate that part of the probability structure of syntax can arise independently of language-specific learning. This identifies human language production as a possible cognitive source of universal statistical structure in language, while providing a data-independent structural bias for probabilistic models, including large language models.

cs.CL↗

Sasaki-Einstein rational homology spheres, rational varieties and the Berglund-Hübsch rule

We find Sasaki-Einstein metrics on rational homology $(4n-1)$-spheres for $n>1$ built from cyclic polynomials of index 1 cutting out rational varieties. The Einstein metrics found here are inequivalent to the ones found by Boyer and Galicki in arXiv:math/0311355. Our findings are consequence of an improvement, for hypersurfaces defined by cycle polynomials, on the estimate given by Johnson and Kollár to determine Kähler-Einstein orbifold metrics. We also construct weighted hypersurfaces that contain the rational varieties described above as codimension two subvarieties and, due to the refined estimate for cyclic polynomials, we find conditions on the weights and degrees of these hypersurfaces so their corresponding smooth links admit Sasaki-Einstein metrics. Finally we study the effect of the Berglund-Hübsch transpose rule on the topology and on the existence of Sasaki-Einstein metrics on the links studied and generalize all the results given in arXiv:2311.15998 for rational homology 7-spheres to rational homology $(4n-1)$-spheres, that is, we show invariance of these two features under the transpose rule.

math.DG↗

Drift Inference for Unit-Root Galton-Watson Processes with Immigration

We study inference on the drift of a critical Galton--Watson process with immigration, a count time series with a unit root. Climate change motivates such nonstationary models for weather-related disaster counts. The drift is the expected increase in the count per period. Ordinary least squares is inconsistent for the drift, so we study state-weighted least squares, giving greater weight to observations following small counts, where conditional variance is lower. Under strict recurrence, we apply null-recurrent regenerative limit theory to obtain a polynomial convergence rate and a standard normal studentized limit. Our main contribution is the recurrence boundary, where we establish a logarithmic convergence rate and a parameter-free non-Gaussian studentized limit. Estimating the optimal weights yields the same first-order limiting distribution as knowing them. Simulations show lower root mean squared error than time-weighted least squares with weights $1/t$, and improved state-weighted confidence-interval coverage when the unit root is imposed.

stat.ME↗

WholeBodyWAM: Learning Whole-Body World Action Models with Scalable Motion Priors

Humanoid whole-body manipulation requires coordinated whole-body dynamics, yet large-scale trajectories from a target robot are expensive to collect and difficult to scale. In contrast, whole-body motion from human and humanoid sources is abundantly available, although such data cannot be directly used as embodiment-specific robot actions. This work asks whether these scalable motion resources can instead provide a transferable predictive prior for humanoid world-action modeling. We introduce WholeBodyWAM, a humanoid world-action model that learns whole-body dynamics from large-scale heterogeneous motion before target-robot training. We curate UniMotion-4K, a motion corpus spanning more than 4K hours from human videos, native 3D motion datasets, and heterogeneous humanoid platforms, and canonicalize these diverse sources into a unified motion space. A language-conditioned Motion Expert is then pretrained to predict future whole-body motion without target-robot action supervision. During robot post-training, the pretrained Motion Expert is integrated with Video and Action Experts through asymmetric Mixture-of-Transformers (MoT) attention, enabling predictive scene dynamics and whole-body motion to jointly inform embodiment-specific action generation. Experiments show that WholeBodyWAM consistently benefits from increased motion-pretraining scale, improves future-motion prediction and downstream task performance, and transfers effectively to real-world humanoid manipulation. Moreover, the pretrained motion prior substantially improves data efficiency under limited target-robot demonstrations.

cs.RO↗

Rank and computation of the pathlifting Jacobian of a DAG ReLU network

This paper provides a self-contained proof of the rank of the pathlifting Jacobian of a DAG ReLU network by performing an induction on the network's number of hidden nodes. In fact, the induction is elementary, and the key recipe is to consider the skeleton matrix of the network, a sparse matrix encoding the network paths, and transform the representation of one of its hidden neurons into an output node. The proof relies on intermediate propositions which link the pathlifting, its Jacobian, the network parameters, and its skeleton matrix, which, on top of permitting to conclude on the rank of the pathlifting Jacobian, also provide a way to compute it without backpropagation and whose computation cost is super efficient in practice compare to usual backpropagation. The paper is provided with a Python module that implements the different propositions of the paper for feed forward networks and is used to experimentally quantifies the computational gain of computing the pathlifting Jacobian with the proposed theory.

stat.ML↗

Total Variation Distance Estimation through Domain Reduction

Computing the total variation (TV) distance between succinctly represented high-dimensional distributions is generally intractable. We give an FPRAS for TV distance between mixtures of product distributions and, more generally, for a natural class of structured probabilistic circuits. Our main technique is a novel application of domain reduction: Given a family of feature vectors indexed by assignments, we use Lewis-weight sampling to replace the assignment domain by a polynomial-size weighted subset that simultaneously approximates the sum of absolute values of every linear projection. For mixtures of product distributions, we construct such reduced domains incrementally over the coordinates, obtaining the first FPRAS with running time polynomial in both the dimension and the number of mixture components. We then extend the approach to smooth, structured-decomposable probabilistic circuits with a common structured architecture.

cs.DS↗

Directions That Don't Drift: Stiefel Manifold Routing for Transformer Attention

The query and key projections $\WQ,\WK$ in attention are almost always trained by Euclidean optimizers with no geometric constraint. We constrain them to the Stiefel manifold and optimize with a Riemannian Adam carrying one scalar second moment per frame---the form of \citet{becigneul2019}, here extended to the compact, non-Hadamard $\St(d,r)$ with a tangent projector, step-norm cap, and polar retraction. Four propositions prove steepest descent in the embedded metric, gradient-scale independence, well-conditioning, and exact $\mathrm{O}(d)$-equivariance. A fifth records that weight decay has \emph{identically zero} Riemannian gradient on $\St(d,r)$ ($W{=}WI_r$ lies in the normal space), so decay cannot act on the constrained frames. On a CIFAR-10 patch benchmark at $n{=}10\mathrm{k}$ this rule gains $\mathbf{+6.79}$\,pp over AdamW across 12 paired starts ($t{=}38.33$, $12/12$); earlier fixed-step Riemannian SGD gains $+1.97$\,pp, of which $+1.69$\,pp comes from frozen orthonormal initialization alone. The corrected Adam's lead grows with data: $+1.9$\,pp at $n{=}1\mathrm{k}$ to $+6.7$\,pp at $n{=}50\mathrm{k}$. A 12-seed ablation credits all gain to the scale-free step ($+4.63$\,pp, $12/12$), nothing to the projector or equivariance; a targeted $\varepsilon$-sweep causally confirms the mechanism ($-2.6$\,pp at $\varepsilon{=}0.1$, $p{<}0.001$). Two five-seed grokking studies confirm the constrained arm does not grok better than the baseline ($p{=}0.019$, A2 wins): the weight-decay exemption has no grokking consequence. A single-seed pilot exploiting this localization achieves the first stable grokking under slingshot conditions---Stiefel + targeted circuit regularization keeps routing-frame isometry error $10^6\times$ lower than the unconstrained ablation through every collapse.

cs.LG↗

Probing submillimeter number counts below the confusion limit: extreme-value statistics of the P(D) distribution and its modulation by gravitational lensing

The shape of the submillimeter galaxy number counts below the confusion limit is a key record of cosmic star formation but is accessible only statistically, through the one-point distribution of map surface brightness, $P(D)$. Classical $P(D)$ analysis compresses the counts into flux-integrated constraints and requires a full instrument forward model. We introduce an extreme-value-theory analysis of the confusion $P(D)$ tail: the peaks-over-threshold formalism, in which exceedances above a threshold $u$ follow a generalized Pareto distribution (GPD). The GPD shape parameter $ξ(u)$ is a flux-resolved, normalization-free readout of the local logarithmic slope of the counts, and its gravitational-lensing modulation $Δξ(u)$ probes their local curvature. We derive analytic relations for both, validate them with end-to-end simulations, and apply the method to Planck 857 GHz maps and the Herschel/SPIRE 350 $μ$m map of GAMA-09. At Planck's 5' resolution the tail reflects the bright, clustered sky rather than the faint counts, though the background alone excludes the single power-law count model. At SPIRE resolution $ξ$ rises markedly with threshold, consistent with the strongly lensed bright population (a first detection of lensing in a $P(D)$ tail), and the bright-masked map favors the Schechter model. Behind galaxy clusters we set the first calibrated upper limits on $Δξ(u)$. CCAT/FYST should separate the count models directly, but a cluster-lensing detection needs more $10^{15}\,M_\odot$ clusters than the sky contains. The GPD tail statistic thus discriminates the functional form of the counts at fluxes of order the threshold, below the detection limit, invariant to map mean, gain and count normalization, and robust to clustering; lensing supplies a calibrated ruler for count features, whose detection awaits deep, high-resolution surveys of massive clusters.

astro-ph.CO↗

Mutual Information as a Tool for Optimal Classification: Application to Identifying Rapid-Responding Behaviour

Existing methods for identifying rapid-responding behaviour in large-scale assessments require parametric assumptions about the population. In this study, we propose a novel, non-parametric, mutual information-based framework of methods as an alternative. The methods within this framework compute the mutual information of the observed responses and discretised response times and maximise the information gain to determine a threshold that differentiates rapid responses from engaged responses. The only difference between the methods is the number of categories in the relevant variables. We present three methods explicitly. The first method uses response correctness and binarised response times. The second method uses correctness and categorises time into three groups. The third method uses raw responses and binarised times. We applied these methods to mathematics achievement data collected through the Programme for International Student Assessment in 2022. Furthermore, we examined the behaviour and usability of the first method in certain realistic conditions at the population and realised levels. The results indicated that the proposed framework is a viable alternative for identifying rapid responses. The framework's novelty lies in its non-parametric nature and its ability to utilise raw responses instead of correctness. Finally, we discuss some possible future research directions on this topic.

stat.ME↗

Placement Is Free, Composition Is Not: The Latin Square as a Provably-Balanced Construction for Heterogeneous Sequence-Mixer Stacks

Since GPT, most Transformers have repeated the same attention mechanism at every layer. Yet this design is largely a convention rather than a tested conclusion. When multiple sequence mixers are combined in one stack, improvements may arise from mechanism choice, placement, or both, making causal attribution difficult. We introduce Aether-7B-5Attn, a 6.59B-parameter mixture-of-experts model ($\approx$2.98B active) whose 49 layers contain seven sequence-mixing mechanisms arranged as a $7\times7$ Latin square. Because each mechanism appears exactly once in every row and column, the design guarantees balanced exposure across depth while eliminating placement confounds. To evaluate this principle, we build a parameter-matched proxy with four mechanisms arranged as a $4\times4$ Latin square over sixteen layers, matched to 700.9M parameters and trained with eight seeds per arm. The results reveal a clear dissociation. Rearranging a distributed heterogeneous stack into a balanced periodic cycle changes validation loss by only 0.16\%, indicating that exact placement has little effect. In contrast, clustering the same mechanisms into contiguous depth bands incurs a 0.59\% penalty, while replacing the heterogeneous stack with a homogeneous one incurs a 1.68\% penalty. These results indicate that performance depends primarily on heterogeneous composition distributed across depth rather than on any particular permutation. We confirm this finding at 2.16$\times$ larger scale (1.514B parameters), where the homogeneous-stack penalty increases to 2.63\% and removing the SSM-family mechanism produces a 3.20\% degradation. We further report per-mechanism cost profiles, English and Korean evaluations, and a causal-safety audit of all 49 layers. We release model weights, training recipes, training code, logs, and architecture source code.

cs.LG↗