Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

P3MaZe: a Mass-Zero constrained-dynamics formulation of particle-mesh electrostatics

We introduce P3MaZe, a real-space particle-mesh electrostatic method that combines the standard short-range/long-range decomposition of Particle-Particle Particle-Mesh (P3M) electrostatics with the Mass-Zero constrained dynamics (MaZe) framework. In this formulation, the smooth long-range electrostatic potential is represented on a mesh as a zero-inertia auxiliary field, while the discretized Poisson equation is enforced as a holonomic constraint during molecular dynamics. By retaining the standard P3M decomposition, P3MaZe preserves the systematic accuracy controls associated with the real-space cutoff, the Ewald splitting, the mesh spacing, and the charge assignment procedure, while replacing the conventional multigrid Poisson solver by a constrained correction problem. The method is validated for molten NaCl and simple point-charge flexible water (SPC/Fw). Values of the potential, structural, translational, collective, and rotational dynamical observables are in quantitative agreement with those obtained with established electrostatic methods, including real-space P3M, and Ewald summation. We further show that, for the linear, time-independent Poisson operator considered here, the constrained correction problem is algebraically equivalent to a direct Poisson solve initialized by a chronological extrapolation of the mesh potential - a connection not previously established for MaZe-based approaches. This result identifies the precise mechanism through which acceleration of the iterative long-range solve is achieved, and we exploit it systematically through higher-order predictors to further improve convergence, while retaining the expected linear scaling with system size. Because the underlying extrapolation is independent of the constrained-dynamics formalism, it can, of course, also be used directly to accelerate conventional iterative real-space Poisson solvers.

physics.comp-ph↗

Fork-Think with Confidence

Parallel thinking has enjoyed great success for boosting LLM performance on reasoning tasks without the need for any re-training. However, existing methods follow a think-first-then-decide paradigm, i.e., they first sample multiple reasoning paths, which inevitably leads to overgeneration, then prune or stop unnecessary paths to compensate. In contrast, decide-first-then-think, i.e., first identifying points that are likely to lead to desirable generations, has been underexplored so far. Following this paradigm, we propose Fork-think with confidence, that first identifies forking points using model confidence in a single seeding path, then triggers thinking, sampling multiple continuations and aggregating them for the final response. Our experiments across three models and three reasoning benchmarks show that Fork-think reduces the token consumption by up to 30% and run-time by up to 57%, while performing comparable to or better than parallel thinking. Our analysis reveals that Fork-think is able to identify forking points that are meaningful with respect to the downstream task and that sampling at later positions can lead to substantially better generations. Finally, we demonstrate how combining Fork-think with existing mechanisms such as early stopping and weighted voting can further boost the performance and perform comparably to existing state-of-the-art methods, without requiring any warm-up or offline training. Our results establish pre-determined forking as a promising research direction for efficient LLM reasoning.

cs.LG↗

A Dominant Self-Conditioning Direction Drives Repetition in Unconditional Continuous Diffusion Language Models

Continuous diffusion language models offer an alternative to autoregressive generation, but their generations may suffer from repetition. We find that unconditional generations from ELF, a recent family of continuous diffusion language models, are more repetitive than human text, while Gen-PPL, a common likelihood-based metric, gives lower perplexity to repetitive generations and can conceal this problem while biasing quality evaluation. Our analysis links this behavior to a self-conditioning feedback loop in which clean-embedding predictions are repeatedly carried into subsequent denoising steps, driving representations toward an effectively one-dimensional contractive attractor associated with repetition. Based on this mechanism, we introduce Attractor-Contrast-Escape (ACE), a training-free inference-time intervention that estimates a repetition direction by contrasting denoising paths trapped in repetition with paths relatively free of repetition and subtracts it from the self-conditioning feedback during denoising. Using a direction estimated only once on ELF-B, ACE reduces mean 4-gram self-repetition rate from 7.28% to 4.48%, while retaining competitive results on several text-quality metrics beyond Gen-PPL. The direction remains effective across ELF sizes and inference configurations, and ACE also generalizes to other unconditional self-conditioned continuous diffusion language models. These results identify self-conditioning feedback as a source of repetition in continuous diffusion language models and show that ACE can directly mitigate this repetition during inference.

cs.CL↗

Elastic deuteron-deuteron scattering in the ${}^5S_2$ channel within Nuclear Lattice Effective Field Theory

We calculate low-energy deuteron-deuteron scattering in the spin-quintet $^{5}S_2$ channel using nuclear lattice effective field theory. The calculation combines chiral interactions at next-to-next-to-next-to-leading order, implemented through wavefunction matching, with the adiabatic projection method. Because the radial cluster basis develops small norm-matrix eigenvalues at large Euclidean projection time, we investigate two stabilization procedures: Tikhonov regularization and projection onto well-resolved norm eigenmodes. The two procedures yield consistent Coulomb-subtracted phase shifts within their statistical and numerical uncertainties. A Coulomb-modified effective-range analysis gives ${}^5a_{dd} = (11.99 \pm 0.22)$ fm and ${}^5r_{dd} = (3.44 \pm 0.79)$ fm. The phase shifts are more negative, and the scattering length is substantially larger than in previous calculations, corresponding to a stronger effective repulsion in the $^{5}S_2$ channel. These results provide a first nuclear-lattice benchmark for deuteron-deuteron scattering and establish a starting point for future coupled-channel calculations of the deuteron-induced reactions relevant to big-bang nucleosynthesis.

nucl-th↗

Generative Refinement for Low-Budget Black-Box Optimization

Black-box optimization is a fundamental tool in science and engineering for optimizing objectives when gradient information is unavailable. It becomes especially difficult when the objective function is expensive to evaluate, limiting the evaluation budget to a few tens or hundreds of queries, and when good solutions occupy complex, low-measure regions of the search space. Generative models can supply useful structural priors in such settings, but existing generative BBO approaches bring significant evaluation cost. We identify three design principles for generative optimization under such low-budget conditions: avoid objective learning, optimize in candidate space, and make every evaluation count. Together, these principles motivate separating structural modeling from objective-driven search. We instantiate them in SPARROW, a simple sequential optimizer that maintains a persistent, ranked archive of evaluated candidates and uses a fixed, unconditional generative sampler solely as a corruption-refinement operator. SPARROW requires only access to the sampler's corruption and refinement processes, and never needs to evaluate the objective to train or guide it. Across three complementary settings, probing thin feasible geometry, disconnected high-performing regions, and failure-prone evaluations, SPARROW outperforms classical and generative baselines under strict evaluation budgets. These results demonstrate that separating structural priors from objective-driven search can be effective when evaluations are scarce and the search geometry is challenging.

cs.LG↗

Novelty Search with Cross-Task Collaborative Discovery

Novelty Search (NS) promotes exploration by rewarding behaviorally novel solutions rather than directly optimizing predefined objectives, making it particularly useful when objective guidance is sparse or deceptive. However, existing NS methods are typically designed for a single search task. When multiple related NS tasks are considered independently, their search processes may repeatedly explore similar regions or rediscover solutions that could be useful across tasks, leading to redundant use of the evaluation budget. To address this issue, we formulate a multitask novelty search setting and propose Multitask Novelty Search with Cross-Task Collaborative Discovery (MTNS-CoD). The central idea is to coordinate discovery across tasks so that different task-specific searches explore complementary regions while useful discoveries can still be reused across tasks. Specifically, MTNS-CoD introduces a multitask repulsion mechanism to discourage redundant exploration in similar genotype regions, together with an adaptive inter-task transfer mechanism that adjusts transfer probabilities according to the observed utility of cross-task exchanges during evolution. Their joint use promotes complementary exploration while selectively reusing beneficial discoveries. We further extend MTNS-CoD to novelty-augmented optimization, where behavioral novelty and objective information are jointly considered to support exploration under deceptive objective landscapes. Experiments on synthetic benchmarks, deceptive maze navigation, MuJoCo policy optimization, and generative novelty search show that MTNS-CoD can improve behavioral discovery and search coverage over the considered single-task and multitask baselines, with additional benefits observed on problems containing deceptive objectives.

cs.NE↗

Deconfined criticality between an antiferromagnetic insulator and a nodal d-wave superconductor: a quantum Monte Carlo study

We present a quantum Monte Carlo study of the transition between the insulating Néel state and the nodal $d$-wave superconductor on the square lattice at half-filling. We access a regime of frustrated magnetic order without a sign problem using a parton representation of the electron in terms of fermionic spinons and bosonic chargons. Both partons move in a background $π$-flux (so the electron experiences no net flux) and are coupled to a quantum fluctuating SU(2) lattice gauge field. In contrast to earlier studies directly on the electronic degrees of freedom, we find evidence for a second-order deconfined quantum phase transition at which both the Néel and $d$-wave superconductivity orders vanish continuously. We compute correlators of the spinon-chargon composite with the same quantum numbers as the electron: we find a gapless Dirac dispersion inside the $d$-wave superconductor, turning into a gapped dispersion in the antiferromagnet.

cond-mat.str-el↗

Uniformly Positive Mean Dimension

We study the relation between uniformly positive entropy and uniformly positive mean dimension at the level of fixed open covers. To a symbolic system X, we associate a hub-and-spoke system Spoke(X), obtained by replacing each symbol by a one-dimensional spoke attached to a common hub. We prove that if X admits a shift-invariant measure of full support, then Spoke(X) has completely positive mean dimension. We also prove that if X has uniformly positive entropy, then Spoke(X) has uniformly positive mean dimension. Finally, using symbolic codings of irrational rotations on tori, we construct hub-and-spoke systems with completely positive mean dimension but without uniformly positive mean dimension or uniformly positive entropy. The examples are nondegenerate: the relevant covers have zero mean dimension and zero entropy, but when refined by iterating under the dynamics the corresponding covering numbers are unbounded.

math.DS↗

Quantum space-depth tradeoffs for coherent block encodings

Block encodings are a basic interface between quantum algorithms and linear algebra. Standard LCU constructions achieve optimal circuit depth but typically require logarithmically many ancilla qubits. We ask how much quantum workspace can be reduced without sacrificing circuit depth, and study this tradeoff from both algorithmic and lower-bound perspectives. For a Hermitian decomposition $A=\sum_{j=1}^L α_j H_j$, with $\|H_j\|=1$ and $α=\sum_j|α_j|$, we give two coherent $\varepsilon$-approximate block-encoding constructions. The first uses one ancilla qubit and has depth $\widetilde O(L(α/\varepsilon)^{o(1)})$, while the second uses $O(\log\log(α/\varepsilon))$ ancillas and achieves depth $\widetilde O(L)$. For a broad Suzuki-based coherent simulation architecture, we prove an ancilla-depth tradeoff. In the polynomial-resource regime and for a constant number of coherent rounds, $\log(1/\varepsilon)\le O((\log Q)^2+2^a\log Q)$, where $Q$ is depth normalized by the number of Hamiltonian terms and $a$ is the ancilla count. Thus polylogarithmic dependence on $1/\varepsilon$ requires more than constantly many ancillas within this architecture. In a separate repeated-query LCU model, for balanced coefficients $1/L$ and error $\varepsilon=η/L$ with fixed $0<η<1$, we prove $2^a=Ω_η(L^2/(T+L))$, where $T$ is the number of oracle queries. Hence $a=Ω(\log L)$ when $T=O(L^α)$ for some $α<2$. Moreover, in the exact case, $a\ge \log L$ regardless of $T$. We also extend this tradeoff to arbitrary nonnegative coefficients. Finally, we apply our low-ancilla constructions to normalized trace estimation in DQC1, obtaining an optimal algorithm linear in the approximate degree together with a matching query lower bound. Together, these results establish quantitative space-depth and space-query tradeoffs in two natural circuit models.

quant-ph↗

Adaptive Reparametrized Time for Score-Based Diffusion Sampling

We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid. Uniform and hand-crafted schedules are standard choices, but they rely on fixed prescriptions and can therefore be suboptimal. To address this limitation, we propose Adaptive Reparameterized Time (ART), a continuous-time control formulation that learns a time change by treating the speed of the sampling clock as the control, so that a uniform grid on the learned clock induces adaptive timesteps in the original diffusion time. Based on a leading-order Euler error surrogate, ART provides a principled objective for allocating timesteps along the sampling trajectory. To solve this deterministic control problem, we introduce ART-RL, an auxiliary randomized formulation with Gaussian policies that turns schedule learning into a continuous-time reinforcement learning problem. We prove that the randomized ART-RL formulation is equivalent to ART at the optimizer level, in the sense that its optimal Gaussian policy recovers the optimal ART time-warping rate through its mean. We further establish policy evaluation and policy improvement characterizations and derive trajectory-based moment identities that yield implementable actor--critic updates for learning the schedule. Across experiments ranging from controlled low-dimensional settings to image generation, ART-RL can be plugged into existing diffusion samplers by changing only the timestep grid, consistently improving sample quality over strong baseline schedules at matched budgets while leaving the rest of the sampling pipeline unchanged. The learned schedules also exhibit broad generalization, transferring without retraining across sampling budgets, datasets, solvers, pipelines, and representation spaces.

cs.LG↗

Bridge-WA: Learning Action-Relevant World Dynamics for Robotic Manipulation

General-purpose VLA models leverage large-scale vision-language priors to understand scenes and instructions, but primarily generate actions directly from the current observation. WAMs further model future scene states, offering a broader view of how the environment may evolve. However, for robotic manipulation, predicting the entire future scene can introduce information beyond what is necessary for action generation; more importantly, effective actions require knowing not only what the scene may become, but also where and how relevant changes unfold. To bridge action generation with these action-relevant aspects of the world, we present Bridge-WA, a general world-action framework that learns complementary representations of future states, spatial changes, and local motion. Specifically, Bridge-WA consists of a Latent World Dynamics Module (LWDM) and WorldBridge. LWDM predicts future states, spatial changes, and local motion from VLM outputs, supervised by corresponding world targets. The WorldBridge are embedded into the action transformer and inject layer-specific combinations of these world priors through multi-source attention, spatiotemporal biases, and reliability-gated feature modulation. This design grounds action generation in action-relevant future dynamics while adaptively regulating world guidance, enabling robust generalization to visual disturbances and viewpoint shifts. We evaluate Bridge-WA on four simulation benchmarks and the real robots, where it achieves maximum success-rate improvements of 11.1%, 42.0%, 3.7%, 23.4%, and 11.1%, respectively, and achieves state-of-the-art average success rates on LIBERO-Dynamic, RoboTwin 2.0 and real-world. In particular, Bridge-WA demonstrates strong generalization to visual variations and viewpoint shifts in both simulation and real-world settings. Code and visualizations are available at: https://hcplab-sysu.github.io/BRIDGE-WA.

cs.RO↗

Exact amplitude relations for diffusion-limited aggregation

It has been known for several decades that the third moment of the multifractal spectrum of the harmonic measure for diffusion-limited aggregates is linked to the underlying fractal dimension of the cluster. We demonstrate, using an argument based on the Hastings-Levitov formulation of diffusion-limited aggregation (DLA) in two dimensions, an even stronger link, connecting the universal amplitude of the third moment to the cluster fractal dimension. This argument can be used for both the standard circular DLA as well as DLA in a cylinder (i.e., with periodic boundary conditions); in the latter case the relationship of the amplitude to the fractal dimension is weaker, and must be extracted via a scaling analysis.

cond-mat.stat-mech↗

PhysMirror: Physics-Aware Mirror Object Generation

Synthesizing physically accurate mirror reflections remains a fundamental challenge for modern text-to-image diffusion models, which are increasingly critical for generating synthetic training data for embodied AI and robotic perception. These models typically struggle with strict geometric constraints, leading to hallucinations that degrade the utility of the synthetic data. To address this, we introduce a novel, end-to-end physics-aware generation framework namely PhysMirror that natively enforces projective geometry through explicit 3D spatial priors. Our method automatically lifts prompted objects into 3D meshes and constructs a lightweight, mathematically exact mirror scene within a simulated environment. By rendering this explicit 3D scene, we extract precise 2D conditioning elements, such as depth maps and segmentation maps, that serve as robust guiding signals for downstream diffusion models, guiding them to generate images with physically correct mirror reflections. Moreover, we introduce Mirror Consistency Score (MCS), reference-free, fully automated metric that quantifies physical correctness using dense feature matching and vanishing point convergence. Experimental results on our newly constructed MirrOB dataset demonstrate that our approach outperforms state-of-the-art baselines in reflection accuracy and physical realism, while maintaining strong text-to-image semantic alignment, providing a reliable pipeline for embodied AI data generation. The source code is released at https://duyphuc0701.github.io/PhysMirror.

cs.CV↗

FedSPM: Routing-Enabled Federated Learning under Dual Heterogeneity via Semiparametric Mixture

Routing-prediction federated learning has emerged as a new paradigm that reframes inter-client heterogeneity as a resource for system-level intelligence: at inference time, the server routes each external query to the best-matched client for prediction. Existing approaches, however, typically treat each client as internally homogeneous, overlooking latent subpopulations within local data. For example, patients with the same diagnosis at one hospital may exhibit morphologically distinct disease subtypes. The coexistence of inter-client and intra-client heterogeneity, which we call dual heterogeneity, can impair both routing and prediction. To address this challenge, we propose FedSPM, a routing-enabled semiparametric mixture framework that represents each client using client-specific latent components. Each component combines a predictive distribution for classification with a feature distribution for routing. To flexibly model feature distributions while effectively sharing information across clients, FedSPM models their density ratios relative to a common nonparametric measure estimated via empirical likelihood. We develop a federated expectation-maximization algorithm that optimizes a tractable surrogate and prove convergence of the exact profiled objective at the standard $\mathcal{O}(1/\sqrt{T})$ rate when the surrogate errors are properly controlled. Experiments on controlled benchmarks and real-world medical data demonstrate consistent improvements in routing and prediction under dual heterogeneity. Code is available at https://github.com/zijianwang0510/FedSPM.

cs.LG↗

XS-VLA: Teaching Tiny Vision-Language-Action Models with Spatial Supervision and Demonstration Conditioning

How can richer training supervision improve robot control while keeping the deployed policy compact? We present XS-VLA, a staged training framework using teacher-derived spatial labels and demonstration-conditioned action learning. Coarse-Grained Spatial Distillation (CSD) initializes the backbone through an auxiliary region-label task. Latent Flow Matching (LFM) then conditions an action-space velocity field on a demonstration latent, using KL regularization while jointly optimizing the backbone and action modules. The deployed policy contains 243.99M parameters and operates without the teacher or posterior encoder. XS-VLA achieves 90.25% average LIBERO success in each of two training seeds, compared with 86.00% for a SmolVLA-256M base trained under our settings. Ablations examine both training stages through matched image pretraining and Huber/MSE controls. On three Mobile ALOHA tasks, average strict success increases from 21.7% to 65.0%. These results demonstrate the control utility of auxiliary representation initialization and regularized demonstration-conditioned flow learning for compact VLA~policies.

cs.RO↗

The nonlocal attraction-repulsion transport equation with power kernels

We study a nonlocal continuity equation on $\mathbb{R}^d$ in which a probability density is driven by the competition between attraction toward a prescribed background measure $ω$ and self-repulsion among particles, governed respectively by the power-law kernels $ψ_a(x) = |x|^{1+a}$ and $ψ_r(x) = |x|^{1+r}$ with exponents $a, r \in [0,1)$. We establish global Lagrangian well-posedness via a squared-radius regularization, obtaining {finite-time $L^\infty$ and moment bounds}, $W^{n,\infty}$ regularity, and uniqueness in the Lagrangian class. When the initial data is compactly supported and attraction dominates ($a > r$, or $a = r$ with $ω(\mathbb{R}^d) > 1$), we prove that the support remains uniformly bounded at all time; a counterexample shows this fails for $a = r > 1$. For the attractive-dominant nonquadratic range $0 \leq r \leq a < 1$, we characterize zero-flux stationary states via a free-boundary problem involving a fractional Laplacian operator, reducing the stationarity condition to a fractional exterior Dirichlet problem. This characterization allows us to exhibit explicit examples of stationary measures in dimensions $d \in \{1,2,3\}$. Numerical particle simulations {are consistent} with the theoretical stationary profiles. Finally, for $0 \le r \le a < 1$, we prove that every global solution whose energy is bounded from below and whose moments are bounded uniformly in time converges, along a sequence of times, to a zero-flux stationary state. Full convergence holds when, in addition, the support remains uniformly bounded (and, for $r = 0$, the density remains uniformly bounded) and the omega-limit set contains a single stationary state.

math.AP↗

Radiation-hydrodynamics of star-disc collisions: From system parameters to outflows and lightcurves

Quasi-periodic eruptions (QPEs) are nuclear transients producing bright, repeating soft X-ray flares superimposed on quiescent emission. A promising interpretation is that they are powered by star-disc collisions, in which a star crosses an accretion disc around a supermassive black hole, drives shocks, and launches dense outflows from which radiation emerges. We present a systematic study of star-disc collisions, linking the physical parameters of the collision to the resulting outflows and emerging bolometric luminosities. We perform three-dimensional local radiation-hydrodynamics simulations, varying the disc surface density and vertical density profiles, stellar velocity and radius, and local collision angle. We focus on the regime where the star remains unperturbed by the collision. We find that variations in stellar velocity and accretion disc surface density leave the bow shock and outflow morphology largely unchanged. Faster stars produce brighter flares, while denser discs mainly increase the flare duration. Increasing the stellar radius increases the momentum of the forward outflow and produces brighter and longer flares. More centrally concentrated discs yield brighter and shorter flares because radiation escapes more efficiently through outer low-density layers. More oblique crossings reduce the momentum and luminosity asymmetry of two outflows, and lengthen the flares. We provide empirical scalings of the peak luminosity and flare duration with the individual system parameters and apply them to GSN 069. The best candidate solutions favour a star with a radius $\sim R_\odot$ on a retrograde orbit, colliding with a dense post-TDE disc with a vertically concentrated density profile. Our findings suggest that specific combinations of system parameters can reproduce characteristic flare amplitudes, durations, duty cycles, and strong-weak flare patterns observed in QPE sources.

astro-ph.HE↗

Prompting Image Generators for Training-free Primitive Shape Abstraction

Compact primitive abstractions represent 3D shapes with a few geometric primitives while preserving recognizable components. Learned methods depend on their training classes, and optimization-based methods split shapes geometrically rather than into parts. We instead reuse the visual part knowledge of pretrained models without task-specific training or fine-tuning. A vision-language model names parts in multi-view renders, and an unmodified image generator paints color-coded part masks. Reprojection and spatial clustering recover 3D instances, and a classical optimizer fits one tapered and bent superquadric per part. With five to eight primitives per object, the abstractions match the Chamfer distance of the strongest learned baseline on HumanPrim, improve on it by 10% on Toys4K, and have the lowest overlap among compact methods, while chair legs, backrest bars and wheels remain separate primitives. Our accuracy also transfers better than theirs to objects outside the learned methods' ShapeNet training classes. Replacing the generated masks with part labels from the 3D segmentation methods P3-SAM or PartField lowers IoU by 7 to 17 points under the same fitter. Further studies relate the remaining volumetric error to part granularity and to parts that the rendered views observe from one side only.

cs.CV↗