Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Lithium Niobate Surface-Acoustic-Wave Resonator for Very-High-Frequency Isolated Power Conversion

Piezoelectric devices offer a promising magnetic-free alternative for high-frequency power electronics, enabling compact energy transfer with inherent galvanic isolation. Here we demonstrate a lithium niobate surface acoustic wave (SAW) resonator that enables isolated power transfer and addresses the challenges in existing piezoelectric devices for power applications, including operation frequencies, power delivery figure of merit (FOM), power handling, and electrical isolation. Our application-optimized 40-MHz SAW resonator, with the two electrically isolated interdigital transducers (IDTs) and acoustic Bragg mirrors, achieves a 25.7 W output power at a less than 15 °C temperature rise with only passive cooling through the back mounting side, and an isolation voltage of 2.17 kV. In DC-AC power conversion, we demonstrate a peak efficiency of 93.8 %, a power delivery FOM of 13.47 mW/V$^2$ and a peak chip power density of 81.7 W/cm$^3$. DC-DC conversion is also demonstrated by connecting a diode rectifier to the output IDT, despite efficiency degrading to 52.7 %. Our results advance SAW devices as one of the core components for next-generation very-high-frequency, miniaturized, high-density, magnetic-free, isolated power converters.

physics.app-ph↗

A uniform degree bound for modular rank-two Nahm sums

Let $F$ be a rank-two symmetrizable Nahm sum, with any symmetrizer $\operatorname{diag}(d_1,d_2)$, a positive-definite rational matrix and rational shifts, and let $K$ be the field generated by its Nahm point; write $m=d_2/d_1$. We prove that if $q^cF$ is modular, then $[K:\mathbb{Q}]\le 24$. Modularity forces the higher coefficients of the logarithmic radial asymptotic expansion to vanish; we study the resulting polynomial equations through their Newton polygons and a pole-order analysis along the determinant conic, and for $m=1$ show that their common components occur only on one explicit family, whose Nahm point is rational. For equal row sums, Bloch-group torsion gives $[K:\mathbb{Q}]\le 2$. For complementary Nahm coordinates, the bound is $2$ when $m\ne 1$ and $4$ when $m=1$, using the first two asymptotic corrections in the latter case. Several steps of the proofs rely on exact computation.

math.NT↗

EpiWorld: Grounding LLM Policy Agents in Epidemiological World Models

Epidemic intervention policies are textual artefacts that human decision-makers interpret, justify, and revise through natural language, making large language models a natural candidate for epidemic policy reasoning. A naive LLM, however, lacks the epidemic dynamics needed to project intervention consequences, the quantitative surveillance signals required to assess severity, and the institutional constraints that define admissible actions. We present EpiWorld, a closed-loop framework that grounds an LLM policy actor in a learned action-conditioned epidemiological world model and a tiered skill library of public-health protocols, surveillance tools, and adaptive lessons accumulated through after-action analysis. Given a candidate intervention, the world model predicts regional epidemic evolution and enables fast counterfactual rollouts that provide feedback for policy selection and refinement. Outcomes of simulated futures are distilled into reusable lessons while protocol constraints remain fixed, allowing the decision process to improve without sacrificing interpretability or controllability. We evaluate both the world model and the end-to-end framework on retrospective COVID-19 and Influenza datasets: the world model achieves the best out-of-distribution Peak-MAE among all forecasting baselines, and the closed-loop framework reduces cumulative hospitalisation by up to 59% across datasets and by an average of ~16% across six LLM backbones, outperforming reinforcement-learning and optimal-control policy baselines.

cs.CL↗

AI-Assisted GPU optimization of the Stochastic Variational Method for Few-Body Boson System

The stochastic variational method (SVM) is one of the most powerful methods to solve quantum few-body systems precisely in various fields, such as nuclear and atomic physics. To the best of our knowledge, no SVM code optimized for GPUs has been reported. We developed a parallel SVM code for few-body clusters of $^4$He atoms and optimized it for two GPU architectures, NVIDIA GH200 and AMD Instinct MI300A, with the entire code written by an AI coding agent. On the algorithmic side, we introduce the secular-equation method with the Gu--Eisenstat prescription for solving the generalized eigenvalue problem within the SVM framework. The main part of the code tuning is to batch the many small matrices so that the GPUs are used efficiently, together with an array layout in memory optimized for coalesced addressing. For the five-body system with 4800 SVM basis states, the tuned SVM code runs 16.2 (14.7) times faster on the NVIDIA GH200 (AMD MI300A) than the same tuned code on the Intel Xeon Max CPU system, and 168 (153) times faster than the first working CPU implementation. The achieved performance brings 6-body and larger cluster systems within reach.

physics.comp-ph↗

Formation and Eruption of a Vortex-driven Magnetic Flux Rope in the Simulated Quiet Sun

Magnetic flux ropes (MFRs) are key structures for understanding flares and coronal mass ejections in active regions, but their characteristics in the quiet Sun remain poorly understood. With a radiative MHD simulation spanning from the upper convection zone to the corona, we analyze the formation and eruption of a supergranular-scale flux rope. We use a clustering method to group the closed field lines by their connectivity. The clusters are gathered further into 8 persistent cluster assemblies (CAs) that correspond to a flux rope and ambient magnetic structures. A persistent counterclockwise vortex is maintained at the converging point of the supergranules, i.e. the positive footpoint of the flux rope. The vortex continuously injects helicity into a magnetic flux tube and wraps ambient magnetic structures around the central flux tube. On a time scale of one hour, the flux tube evolves to a strongly twisted flux rope and erupts. The eruption exhibits similar current sheet structures as inferred from the flux rope eruption in active regions; however, the mass ejection and heating are much weaker and give rise to very insignificant observable features. The synthetic EUV images show mostly dimming features caused by density rarefaction in the expanding flux rope. This work helps elucidate the dynamic and complex evolution of flux ropes formed in the quiet Sun and suggests that similar events in the real Sun may have been underestimated due to their stealth behavior.

astro-ph.SR↗

Lower regularity well-posedness for a higher-order Schrödinger equation with cubic nonlinearities on the half-line

In this paper, we continue the study of Himonas and Yan\cite{himonas2024schrodinger,himonas2026higher,himonas2024higher} on the nonlinear Schrödinger equation with a dispersion of order $2m$ and cubic nonlinearities, \[ iu_t+\left( -1 \right) ^{m+1}\partial _{x}^{2m}u=N_k\left( u,u,u \right), \] where, $m\ge1$ being an integer, \[N_0\left( u,u,u \right) =uuu,\quad N_1\left( u,u,u \right) =\bar{u}uu,\quad N_2\left( u,u,u \right) =\bar{u}\bar{u}u,\quad N_3\left( u,u,u \right) =\bar{u}\bar{u}\bar{u},\] We derive trilinear estimates at lower regularity and thereby prove that the cNLS-2m on the half-line is well-posed at the optimal regularity $s=-\fr{m-1}{2}$. This improves the previous result, $s>-\fr{m-1}{2}$, and answers an open question left in \cite{himonas2026higher}. In addition, for $k=0,2$, we establish well-posedness at $s=-\fr{m-1}{2}$ as well. Moreover, for $k=3$, the nonlinearity exhibits a stronger resonance relation, which implies well-posedness for $s>-\fr{2m-1}{3}$. Our derivation of the trilinear estimates relies on varieties of Strichartz estimates, which differs from the $[k;Z]$-multiplier norm method employed in \cite{himonas2026higher}.

math.AP↗

Initial and initial-boundary value problems for a cubic sixth-order Boussinesq equation

In this paper, we consider the sixth-order Boussinesq equation with a cubic nonlinearity, \[ u_{tt}-u_{xx}+k u_{xxxx}-u_{xxxxxx}+\left( u^3 \right) _{xx}=0, \ \ \mbox{with }k=\pm 1,\] posed on both $\mathbb{R}^+$ and $\mathbb{R}$. The local wellposedness are improved in related Bourgain-type spaces, $X^{s,b}$, for $s>-\frac12$ with $b=\frac12$ and $b>\frac12$ accordingly (See. \cite{esfahani2025well,zhong2024local}). The key ingredient is to refine the trilinear estimate on Burgain type space by adapting ideas from Tao's multiplier method, interpolation theory and our early work on IBVPs of dispersive equations \cite{li2020low,tao2001multilinear}.

math.AP↗

Learning Query Encoders Can Be Hard Even When Vector Retrieval Is Geometrically Easy

Efficient vector retrieval requires both a corpus geometry that supports retrieving the right documents through vector similarity, and a query encoder that can embed queries near their desired documents in the embedding space. Recent work has studied geometric capacity through the lens of the minimum embedding dimension needed to realize all top-$k$ answer sets of $n$ documents. We study a different notion of geometric capacity--the maximum recall achievable for a frozen document index--and explore whether learned query encoders can reach this ceiling. On several real-world retrieval benchmarks, we show that retrieval quality of single-vector query encoders often lies far below what the document indices can support. Motivated by this observation, we give theoretical evidence that learning query encoders can be computationally hard. In particular, we construct a retrieval task that (1) admits a query encoder with perfect recall which is representable by a small one-hidden-layer ReLU network, but (2) any statistical-query learner (a class capturing learners that access training data through aggregate statistics) provably requires exponentially many statistical queries to achieve non-trivial recall advantage over the random baseline $k/n$. Taken together, our results suggest substantial unrealized geometric capacity in retrieval benchmarks and establish query encoder learnability as a possible barrier in embedding-based retrieval.

cs.IR↗

CSIR: Contextually and Socially Informed Robots for Efficient Person Goal Navigation

We address the Person Goal Navigation (PersonNav) problem, enabling robots to search for people under realistic constraints on information accessibility and language uncertainty. This tackles the current solutions revolving around rigid person finding systems requiring exact knowledge of the individual for item-delivery in indoor settings. While the focus in literature is on learned methods using unavailable public or sparse data due to human privacy. We propose a planning framework that combines distance and semantic information about the recipient (habits and intent) weighted by the trust of the user-provided information. A synthetic benchmark of scenarios is also developed including actors, items, and requests to evaluate performance before real-world deployment aiming to simulate natural language human-robot interactions. Results show our informed search outperforms classical distance-based graph baselines, while semantics alone lead to ungrounded, sporadic search. Our method achieves strong gains over baselines, with an LLM-based variant performing comparably, and its explicit belief representation naturally supports future Bayesian filtering. Hardware tests demonstrate our method supports real-world embodiment able to leverage between semantic and distance information in a real setting. This work moves toward more intelligent mobile service agents capable of human-like, informed search in realistic environments to be leveraged in day-to-day use. https://anonymous.4open.science/r/personnavsite-4ED2/index.html.

cs.RO↗

Sharp Liouville thresholds and endpoint rigidity for finite Morse index solutions of the $p$-Laplace Lane--Emden equation

We study $-Δ_p u = |u|^{q-1}u$ in $\mathbb{R}^N$ with $N>p\ge 2$, for solutions that are stable outside a compact set, with no assumption of sign, boundedness or symmetry. Damascelli, Farina, Sciunzi and Valdinoci settled the subcritical range for $p>2$, and treated the supercritical range $p^*-1 q_c(N,p)$ we exhibit positive bounded radial solutions stable outside a compact set. At the Sobolev endpoint, every $C^1$ solution stable outside a compact set has finite energy; for $p>2$ this removes the a priori $\mathcal D^{1,p}$ assumption from the low-Morse-index classification of Farina, Mercuri and Willem and answers their question on the Morse index of the Aubin--Talenti extremals: it is one. At the upper endpoint, $q_c(N,p)$ is the exact threshold for a nonzero homogeneous weak solution on $\mathbb R^N\setminus\{0\}$ to be stable outside a ball, and at the threshold there are exactly two.

math.AP↗

GAANet: Global-guided Asymmetric Attention Network for Audio-Visual Speech Separation

Multi-scale design is crucial for efficient audio-visual speech separation, yet effectively modeling multi-scale information for audio-visual feature fusion remains challenging. We argue that the limited capacity of existing approaches primarily arises from: 1) treating features from different modalities in the same manner, and 2) overlooking the role of global features. To address these issues, we propose a Global-guided Asymmetric Attention Network (GAANet). Our model introduces two core innovations: first, an asymmetric multi-scale fusion framework that allows audio and visual streams to extract and interact with features at their respective optimal temporal resolutions, removing the need for symmetric temporal downsampling; second, a global-guided attention mechanism that compresses each modality into a compact global token with a temporal dimension of one, which then provides high-level semantic cues to guide both intra- and inter-modal fusion across scales. Experiments on LRS2 and VoxCeleb2 demonstrate that GAANet achieves state-of-the-art performance, reaching 16.5 dB SI-SNRi on LRS2 and 14.0 dB on VoxCeleb2, while maintaining a lightweight computational profile with only 3.3M parameters and 19.8G MACs. These results highlight the strong potential of asymmetric temporal modeling and global guidance for efficient and robust multimodal fusion. The source code is publicly accessible at https://github.com/redizzy/GAANet

cs.SD↗

Correcting Guided Diffusion Trajectories with Spectral Alignment

The practical success of conditional image generation hinges on fine-grained differences in condition alignment and visual fidelity. Classifier-free guidance (CFG) is central to this success, but its lack of an explicit criterion makes it difficult to assess whether the guided trajectory is progressing as intended. To address this gap, we show that spectral alignment provides a principled criterion for understanding guidance behavior and improving guided diffusion sampling through adaptive correction. Our analysis identifies the spectra of intermediate states as an indicator of consistency with the expected spectral evolution of the forward process. Based on this observation, we introduce Spectral Correction Guidance, a method that corrects deviations from an analytic reference spectrum during sampling. The proposed method is training-free and applicable across diffusion backbones and conditional generation tasks without modifying the underlying model. Experiments demonstrate consistent gains in preference-based metrics over baseline guidance methods in text-to-image generation and improved generation quality over CFG on ImageNet. These improvements persist across a range of guidance scales and with fewer denoising steps. Our analyses and ablations provide insight into guidance behavior and how the proposed method affects generation quality.

cs.CV↗

Jumping up and down: Denoiser diffusion models for discrete ordinal data

Diffusion models are highly developed in continuous spaces for image and video domains. Recently, major advances have been made for discrete diffusion models for categorical data, specifically in the language domain. In contrast, diffusion models for discrete integer-valued data are less developed, despite the prevalence of this modality, ranging from images and music to gene counts. We introduce Jumping Up and Down (JUD)---a new family of denoiser-based diffusion models for discrete ordinal data. This is the first family of diffusion models for ordinal data which centers around training denoisers, which at the same time allows for bi-directional (up and down) perturbations of the data. The simplicity of the training objective, combined with the flexibility of bi-directional perturbations, leads us to obtain competitive results across different data modalities.

cs.LG↗

FiberGeoText: A Vision-Language Model for Population- Level Organization of Superficial White Matter

The superficial white matter (SWM), a critical brain region for cognition across the lifespan and brain disease, contains abundant short-range association fibers whose organization remains incompletely characterized, in part because the short trajectories and highly variable cortical folding make correspondence across individuals challenging. Anatomically corresponding connections may vary in spatial location across individuals and therefore may not be adequately defined by geometric proximity alone. We introduce FiberGeoText (FGT), a vision-language model (VLM) for organizing short-range superficial white matter (SWM) streamlines reconstructed from ultra-high-resolution diffusion MRI into population-level clusters. FGT jointly represents three complementary properties of each streamline: its three-dimensional trajectory, its cortical anatomical context, and its shape. Cortical endpoint information from multiple parcellation schemes is expressed as text and encoded using a pretrained large language model (LLM), enabling heterogeneous anatomical descriptions to contribute to a common continuous representation. We evaluated FGT on acquired submillimeter 0.76 mm diffusion MRI data. Compared with state-of-the-art (SOTA) methods, FGT produced substantially greater cortical parcel coherence, within-cluster shape consistency, cluster-size consistency, and cross-subject correspondence. The trained model also generalizes well to unseen subjects with an average of 96.7% of the 5,000 learned clusters recovered, and high consistency of cluster structure between training and testing data. Together, these findings demonstrate that integrating geometric, anatomical, and shape information by learning multimodal deep embeddings with a VLM model enables robust learning of population-consistent SWM organization despite interindividual anatomical variability.

cs.CV↗

Motzkin Numbers Count 2-Stack-Sortable Permutations Ending in Their Least Entry

We prove the following conjecture of Zhang (arXiv:2604.10779, Conjecture 6.1): for $n \geq 0$, the number of $2$-stack-sortable permutations of $\{0,1,\dots,n\}$ ending in $0$ is the $n$th Motzkin number. By Zhang's result, there is a bijection between $2$-stack-sortable permutations ending in their least element and standard composition tableaux of width at most $2$. We then show bijectively that there are an equal number of these and standard Young tableaux of width at most $3$, which are known to be counted by the Motzkin numbers.

math.CO↗

Design Automation for Gray-Code Quantum Read-Only Memory

We propose a programmable quantum read-only memory based on Gray code encoding, termed GQROM, designed to efficiently load classical data into quantum circuits. This architecture mitigates critical bottlenecks in near-term quantum computing by reducing initialization overhead and gate-induced errors. By utilizing Gray code, in which consecutive addresses differ by only a single bit, GQROM minimizes state transition complexity. We demonstrate the architecture's flexibility in accessing targeted data through initial-state modification. The entire design, including a systematic EDA flow, was implemented and validated on the Qiskit platform, confirming its practical feasibility. These results underscore GQROM's potential as a scalable and hardware-friendly solution for efficient state preparation in quantum circuits.

quant-ph↗

Absence of universal backscattering immunity in reciprocal photonic waveguides

Suppressing parasitic back-reflection caused by imperfections, such as random defects, material inhomogeneities, localized inclusions, and fabrication-induced geometric deviations, has long been a central objective in optical-waveguide engineering. In this Letter, we theoretically establish that universal immunity to backscattering from arbitrary disorder is fundamentally impossible in reciprocal optical waveguides, irrespective of whether the underlying structure is topologically trivial or nontrivial.We establish this result through two complementary constructions of physically admissible reciprocal perturbations. The first considers weak permittivity perturbations of finite spatial extent within the first Born approximation, whereas the second considers optically small dielectric perturbations of finite contrast. In each construction, an admissible realization exists that produces a nonzero leading-order backward-scattering amplitude. We trace the fundamental origin of this limitation to the absence in photons of the intrinsic spin-$\tfrac{1}{2}$ and fermionic time-reversal structure underlying Kramers protection in electronic systems. Consequently, any mechanism that enforces photonic backscattering suppression throughout a prescribed class of disorder must instead be engineered into the underlying electromagnetic structure and its constitutive response, which can itself be modified by generic reciprocal perturbations. Numerical examples in representative reciprocal topological waveguides further illustrate these limitations for random defects and sharp bends. We further identify several constructive strategies for suppressing back-reflection within certain prescribed classes of disorder.

physics.optics↗

Proprioceptive Sketches as Long-Horizon Intent for Generative Action Policies

Generative robot policies predict short action chunks but lack explicit long-horizon intent. Recent methods expose longer-horizon structure through language plans, subgoal images, or video forecasts, which are costly to generate and still need to be translated into robot motion. Predicting future robot motions avoids this translation, but a dense, time-indexed trajectory requires numerous parameters to cover the full remaining task, and over a short horizon it largely repeats the action chunk and adds little guidance for action generation. We propose Proprioceptive Action Models (PAM), which jointly generate a compact, timing-free sketch of the robot's remaining joint-space path and a dense executable action chunk within a single transformer denoiser. The sketch parameterizes the path by arc length rather than time, capturing geometric intent invariant to execution timing. Block-causal attention and a staggered denoising schedule maintain directed sketch-to-action dependence, ensuring the action tokens condition on a progressively cleaner sketch throughout sampling. In simulation, PAM improves over its action-only counterparts on Push-T and LIBERO-Long; on four real-world bimanual tasks, it raises success from 47.5% to 75.0%. Project page: https://nicehiro.github.io/pam_dp/

cs.RO↗