Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Antiperiodicity in the Duffing--Holmes oscillator: symmetry origin, parity selection, and transitions

We investigate the origin and distribution of antiperiodicity --- oscillations satisfying $x(t+T)=-x(t)$ --- in the periodically driven Duffing--Holmes oscillator, combining analytical arguments with extensive numerical exploration. Antiperiodic orbits are precisely the periodic orbits invariant under the half-period shift symmetry $S:(x,\dot{x},t)\mapsto(-x,-\dot{x},\,t+T_d/2)$ of the equations of motion, with $T_d$ the driving period. We map the antiperiodic regions across the plane spanned by the amplitude and the frequency of the forcing, together with the periodic and chaotic domains and the potential wells visited by each orbit. The invariance under $S$ imposes a parity selection rule, verified without exception across our parameter sweeps: antiperiodic orbits lock to the drive only at odd multiples of the forcing period. Periodic orbits that lack the antisymmetry occur instead as conjugate pairs related by $S$, each orbit being the point reflection of its twin. We further show that antiperiodic orbits cannot bifurcate through a direct period doubling: the symmetry must break first, in a supercritical pitchfork in which the antiperiodic orbit splits into two conjugate, symmetry-broken orbits; alternatively, the antiperiodic orbit disappears with its symmetry intact, in a saddle-node bifurcation with an antiperiodic saddle. Antiperiodicity thus emerges as the orbit-level manifestation of a discrete symmetry of the driven system.

nlin.CD↗

Multilayer-Dynamic Network Clustering with Application to World Trade Data

International trade data, such as the FAO dataset, can be naturally represented as \emph{multilayer-dynamic networks}, where countries are nodes, trade relationships are edges, different products correspond to different layers, and networks evolve over time. An important problem is how to identify evolving community structures in such multilayer-dynamic trade networks. Motivated by this problem, we study community detection in multilayer-dynamic networks, allowing the community structure to vary across both layers and time. We propose a novel method, \emph{MuDySC} (Multilayer-Dynamic Spectral Clustering), which smooths the eigenspace projection matrices across adjacent time points and across layers at the same time point. We develop an efficient alternating iterative algorithm and establish both global and local convergence results, with the latter allowing weaker conditions on the tuning parameters when the relevant eigenspaces are sufficiently close and the algorithm is suitably initialized. We apply MuDySC to the FAO data. The analysis reveals clear asymmetry between export and import community structures and highlights both persistent and shifting trade positions of major countries. As an extension to accommodate substantial heterogeneity across network layers, we develop TLC-MuDySC, a tensor-based layer-clustering method that first identifies structurally related layers and then applies MuDySC within the estimated layer groups.

stat.AP↗

Schwarzschild black holes from twistor space

Twistor theory forms the basis for many surprising advances in areas ranging from dynamical systems to quantum field theory. Yet for almost fifty years, one of the main drawbacks of twistor theory has been its inability to give non-perturbative descriptions of non-chiral (or non-self-dual) field configurations. This difficulty is known as `the googly problem.' In this paper, we provide a resolution of the googly problem for a particular solution of the vacuum Einstein equations: the Schwarzschild metric. We start with the twistor space of the self-dual Taub-NUT Euclidean gravitational instanton, expressed in Kerr-Schild form. Within this twistor space, we then consider a quadric which corresponds to the anti-self-dual Taub-NUT metric. While the full quadric is not holomorphic with respect to the complex structure of the self-dual Taub-NUT twistor space, its holomorphic locus still has complex dimension two. This `coincidence locus' -- points in twistor space on the holomorphic portion of the quadric -- inherits a complex structure from the twistor space and a symplectic form from the quadric itself. Remarkably, these structures are compatible, giving rise to a non-self-dual, four-dimensional Kähler metric which is conformal to Schwarzschild (in Lorentzian or Euclidean signature). This is the first instance of a non-self-dual Einstein metric constructed entirely from holomorphic data in a twistor space.

hep-th↗

Reward Valuation in Large Language Models: Causal Induction of Anhedonia

Recent frontier models mimic complex aspects of human cognition. Here we ask whether this alignment extends into reward valuation, which we assess in a mechanistic framework. Specifically, we use clinical tests that were developed to evaluate anhedonia in human subjects with major depressive disorders. Mechanistically, anhedonia is frequently associated with dysregulation in the Nucleus Accumbens (NAc) and the broader dopaminergic reward system. While neuroimaging has localized these deficits, establishing a causal link between NAc activity and specific behavioral symptoms remains a challenge. We use these ideas from neuroscience to functionally identify reward-anticipatory units in state-of-the-art AI models, and evaluate their causal involvement via targeted perturbations. We find that not only are such model units predictive of NAc brain recordings, their perturbation also induces behavioral effects mirroring human anhedonia: the model opts for low-effort, low-reward tasks in effort-based decision-making paradigms. Crucially, our results demonstrate that this represents a specific deficit in self-centered reward valuation and anticipation--rather than a loss of task capability, reward calculation, or effort avoidance. This induced vulnerability aligns with clinical measures of anhedonia and motivation in humans, such as DARS and MAP-SR, instruments that contain no reward-related vocabulary, ruling out a purely lexical account of the perturbation effect. Taken together, our results suggest reward valuation circuits in AI models that functionally mimic those in humans.

cs.LG↗

Identifiability and likelihood estimation for homogeneous angular regression with random coefficients

Repeated directional responses may vary between individuals in both overall orientation and sensitivity to explanatory cues. We study homogeneous angular regression with correlated Gaussian cue coefficients and a circular random intercept, separating observation concentration from resultant length. Constructive identification conditions recover the joint coefficient law, intercept concentration and observation-error parameters, including unknown uniform contamination. The argument combines repeated contrasts, an arc-support restriction and joint Fourier deconvolution, and extends to exogenous continuous designs. Integration in coefficient coordinates yields smooth population scores even when resultant cancellation prevents a regular latent mode. A fixed-proposal adaptive importance algorithm evaluates the marginal likelihood and its derivatives. In 8,000 independent simulated datasets, 7,998 fits pass numerical checks, and 12,799 of 12,800 local profile intervals are available. Correlation profile-set coverage nevertheless ranges from 91.9% to 94.8% at the nominal 95% level. Greater precision and numerical availability therefore do not ensure calibrated inference. An exploratory extension to 150 Canadian weather stations is evaluated on 259,807 three-hour forecasts from a new 2026 period after fitting on 2023-2025 data. With circular station rotations in all matched models, independent station slopes give modest predictive gains over slopes shared within regions, in all five regions. Estimating their correlations adds no useful predictive gain; this does not assess parameter uncertainty, and predictive calibration remains imperfect. The application uses a noncancelling mean and Laplace integration, so it illustrates the model class rather than testing the proposed algorithm.

stat.ME↗

Stochastic Inversion of Multivariate Uniform-Distribution-Preserving Transformations

A multivariate transformation of the unit cube with component transformations that are piecewise continuously differentiable and uniform distribution preserving (udp) is considered. A stochastic inverse transformation is defined using randomization to overcome the non-injective nature of the udp transformations. The inverse transformation preserves the uniform margins of a random vector distributed according to a copula and yields different copulas for different randomizations. A copula density transformation result for the multivariate stochastic inverse is proved and illustrated in the bivariate case.

math.ST↗

GeoFP8: Geometry-Aware FP8 Gradient Compression for Distributed LLM Training

Gradient communication is a primary scaling bottleneck in large language model (LLM) pretraining. Communicating gradients in low-precision formats, such as FP8 and NVFP4, can significantly reduce the communication volume. Existing methods quantize gradients via linear or nonlinear mappings in Euclidean space, often degrading model performance because highly anisotropic gradients incur direction-dependent distortion. We present GeoFP8, a geometry-informed gradient scaling method that performs low-precision communication in geometry-aware coordinates. By transforming gradients into a near-isotropic space before quantization, GeoFP8 makes low-precision representations substantially more faithful to their high-precision counterparts. GeoFP8 only changes the coordinate system used for low-precision gradient communication and does not change the optimizer, training recipe, communication collective, or low-precision format. We also develop a simplified geometry-aware transformation algorithm with low-rank approximation and selective application to balance the computation overhead and communication reduction. We examine the empirical convergence of GeoFP8 using Llama-300M and Llama-600M models. Our results show that GeoFP8 reduces the end-to-end pretraining time of Llama-600M by 7.6% on 64 NVIDIA GH200 Superchips, while improving the downstream task preservation profile over direct Euclidean FP8 communication under the same optimizer and communication path.

cs.DC↗

Manifold-adapted radial basis functions for reduced-order modelling of chaotic flows

Chaotic systems often evolve on a low-dimensional attractor whose geometry varies from one region to another. We propose a non-intrusive reduced-order model that reads this local geometry by clustering and uses it to shape a radial basis library whose kernels adapt to each region. Fitting the reduced velocity onto this library by one global least-squares solve gives an explicit, differentiable vector field that reproduces the long-term statistics without any use of the governing equations. A radial basis field decays away from the data and cannot by itself return an escaped state. The integration is therefore stabilised by a kinematic corrector, whose reported magnitude measures how far each result rests on the learned field. On Lorenz-63 the model recovers the attractor, its marginal densities and its Lyapunov spectrum. On Lorenz-96 its valid prediction time matches typical configurations of neural-network and reservoir-computing forecasters and trails their best-tuned ones, and the invariant measure is reproduced on both the full state and on a reduced observable. On the Kuramoto--Sivashinsky equation and the quasiperiodic Kolmogorov flow the model matches the energy distribution and spectrum of an intrusive quantised-local Galerkin model and improves on a global Galerkin projection of the same reduced dimension. The recovery is dictated by the distance from a state to its nearest kernels, not by the one-step regression error.

physics.flu-dyn↗

A note on Zamolodchikov's recursion relation for the torus conformal block and its light limit

In this paper, we review the Zamolodchikov-like recursion relations for torus one-point conformal blocks in both Liouville and $A_2$ Toda theories. Starting from these relations, we derive the corresponding recursion relations in the light asymptotic limit. In Liouville theory, the light-limit recursion reproduces the known expression for the one-point light conformal block in terms of the Gauss hypergeometric function. For $A_2$ Toda theory, our recursion relation provides a new, efficient method for explicit calculations, and we have verified that it is in full agreement with previously known results in the literature. Last but not least, we present closed-form expressions for the coefficients of the insertion-point expansion for both the generic and light $A_2$ Toda one-point blocks.

hep-th↗

Simon-SR: Spatially Adaptive Modulation and Visual Prompt Adaptation for Text-Reinforced Super-Resolution

Severe downsampling makes single image super-resolution (SISR) an ill-posed problem, in which the language-guided multi-modal methods are especially vulnerable to erroneous priors and require costly textual annotations. To address these issues, we propose Simon-SR, a multi-modal SISR framework leveraging learnable prompts for efficient semantic mining and robust text-image fusion. Simon-SR treats textual semantics as learnable latent variables. Specifically, the Contrastive Prompt Learning (CPL) mines instance-level semantics from unannotated images with frozen CLIP encoders, and Prompt-Guided Spatially Adaptive Refinement (PSAR) injects them through attention-gated multi-modal fusion. On CUB and COCO2017 at $\times4$ and $\times16$, Simon-SR gains up to 0.50 dB PSNR and 0.0133 SSIM over current SOTA. The project is available at https://github.com/CHT05017/ICAIS26-SimonSR

cs.CV↗

A solvable normal form for coupled swarmalators

Swarmalators are mobile generalizations of phase oscillators. Introduced to model systems in which sync and self-assembly interact, they remain poorly understood theoretically. Unlike the Kuramoto model for coupled oscillators, existing swarmalator models lack a normal-form foundation, and their basic stabilities and bifurcations remain largely unsolved. Here we address both problems. Building on Tanaka's reduction of chemotactic oscillators, we show that the canonical one-dimensional swarmalator model -- previously introduced as an ad hoc toy model -- is recovered in the first-harmonic, zero-lag limit, implying its behavior is generic. We then derive the stability boundaries organizing its four collective states, show they meet at a single cusp, correct a previously published order-parameter formula, and uncover a non-monotonic sync response absent in the Kuramoto model.

nlin.AO↗

A Koszul complex in quaternionic analysis and its applications

Let $n\geqslant 1, Ω\subset\mathbb{H}^n $ be a domain. We construct a Koszul-type complex for the ideal sheaf $\mathcal{I}_X^{(k)}$ of $k$-regular functions vanishing on $X=\{(q_0, q_1, \cdots, q_{n-1})\in Ω: q_0=0\}$ in several quaternionic variables: $$0\to \mathcal{R}^{(k+2)}\xrightarrow{\widetilde{\mathscr{L}}^{(k)}} \mathcal{R}^{(k+1)}\oplus\mathcal{R}^{(k+1)}\xrightarrow{\mathscr{L}^{(k)}} \mathcal{I}_X^{(k)}\to 0,$$ where $k\geqslant 0$, $\mathcal{R}^{(k)}$ is the sheaf of $k$-regular functions on $Ω$, $\widetilde{\mathscr{L}}^{(k)}=(-L_1^{(k+2)},L_0^{(k+2)})^{T}$, $\mathscr{L}^{(k)}=(L_0^{(k+1)},L_1^{(k+1)})$, and $L_0^{(k)},L_1^{(k)}$ are multiplication-like operators on $k$-regular functions. This gives the quaternionic analogue of the classical Koszul complex. And we present the long exact sequence in cohomology for the case $Ω\cap\{q_0=0\}=\emptyset$ with explicit differential connecting maps, by applying the Cauchy-Fueter complex and cohomological methods. As an application, in the special case $n=1, k=1$, the operator pair $(L_0^{(1)}, L_1^{(1)})$ is shown to be surjective if and only if $H^3(Ω, \mathbb{R})=0$. Furthermore, a cohomological vanishing criterion is given for $H^1(Ω,\mathcal{I}_X^{(k)})$; under this criterion, every $k$-regular function on $\{q_0=0\}\capΩ$ extends to a $k$-regular function on $Ω$.

math.CV↗

SETA: Scaling Environments for Terminal Agents

Large language models (LLMs) are rapidly shifting toward agents that solve tasks through diverse interfaces, including web and graphical user interfaces (GUIs). Among these, the terminal command line provides a text-based, general-purpose interface, covering tasks from system operations to data science and machine learning. However, scaling terminal-agent training remains challenging, as it requires diverse and coherent task instructions, executable environments, and reliable verification, while lacking naturally grounded supervision data. In this work, we propose SETA, a scalable framework for generating verifiable terminal environments for reinforcement learning (RL). The framework consists of two pipelines sharing a unified verification mechanism: SETA-Synth converts diverse sources into standardized RL environments, and SETA-Evol further expands from existing environments with adaptive control of difficulty and diversity. Together, we construct and release SETA-Env, the largest open-source verifiable terminal RL dataset to date, containing over 4,500 environments. We evaluate our dataset by training Qwen3-8B with GRPO on SETA-Env, achieving 12% pass rate on Terminal-Bench 2.0, the best reported result for an RL-trained model at the 8B scale. We further observe gains on DeepSeek-V4-Flash under the same terminal agent harness, with pass@1 on Terminal-Bench 2.0 improving from 40% to 43% and pass@5 improving from 54% to 58%. These results demonstrate that SETA- Env provides high-quality training environments for terminal agents and serves as a valuable resource for advancing research on terminal-based agent learning.

cs.AI↗

On Small Doubling in Right-Ordered Groups and Baumslag-Solitar Groups-II

Recently, Mohan et al. [Results Math. 80 (2025), No. 4, 122] answered Freiman's $3k-4$ conjecture in right-ordered groups under certain restrictions. In this paper, we take a step further by investigating the structure of nonempty subsets $S$ of a right-ordered group satisfying the small doubling condition $|S^2| = 3|S|-3$. Moreover, we provide a complete characterization of all nonempty finite subsets $S$ of the Baumslag-Solitar group $\mathrm{BS}(1,q)$ (with $q \in \mathbb{Z}$ and $q \neq -1$) for which $|S^2| = 3|S|-3$ and the identity element is the minimum of $S$.

math.NT↗

The combinatorics of sector renormalization

The goal of this note is to systematically develop the fundamental arithmetic and combinatorial properties of the sector renormalization operation on rigid rotations. We employ the specific framework of modified continued fractions appropriate for sector renormalization and analyze their properties. By allowing infinite first return times, this framework yields a dynamical compactification of the space of irrationals called the parabolic compactification; we show that it is characterized by some universal properties. We also discuss the corresponding natural extension and introduce the continuant group, the time semigroup, and topological cascades. For example, we demonstrate how a bi-infinite tower of sector renormalizations of irrational rotations can be packaged within a single dynamical plane as a cascade of translations. This note will serve as a foundational combinatorial tool for studying the geometric properties of sector renormalizations of holomorphic maps with irrationally indifferent fixed points, particularly neutral quadratic polynomials.

math.DS↗

Advancing Optimal Subset Oracle via Learning Relaxation of Neural Set Functions

Learning neural set functions is pivotal to a wide range of important applications, including compound selection in AI-driven drug discovery and product recommendation. Recent work has introduced optimal subset oracles to implicitly learn set functions under practical weakly supervised settings, where model parameters are optimized through mean-field variational inference. However, these frameworks rely on Monte Carlo sampling to estimate gradients of the evidence lower bound when updating the variational distribution. Repeated sampling across iterations incurs substantial computational overhead, while the resulting stochasticity can destabilize the optimization trajectory. In this work, we reinterpret the evidence lower bound as a continuous relaxation of the set function and learn a surrogate objective that replaces sampling-based ELBO gradient estimation during variational optimization. The learned surrogate provides stable and efficient gradients throughout the continuous domain, thereby reducing computational overhead and accelerating inference. Furthermore, we establish an approximation guarantee for the proposed framework under submodular maximization and characterize its connection to variational free energy. Experiments on a variety of real-world tasks demonstrate consistent improvements over existing baselines.

cs.LG↗

Mitigation of Initial Transients in Total-f Gyrokinetic Turbulence Simulations Using Neoclassically Relaxed Distribution Function

Total-f five-dimensional gyrokinetic simulations are essential for self-consistent studies of multiscale, multiphysics transport in the edge region of diverted tokamak plasmas. However, conventional initialization with a local Maxwellian distribution often generates large-amplitude transients, particularly geodesic acoustic modes (GAMs). These transients are especially severe in the plasma edge because of steep profile gradients, strong radial electric fields, and high safety factors, and they can interfere with early-time turbulence and transport diagnostics. To address this problem, we present a new initialization scheme for the total-f XGC code that uses a relaxed particle distribution obtained from a computationally inexpensive axisymmetric simulation. Before the distribution is transferred to the full turbulence simulation, phase-space smoothing is applied to reduce particle noise while preserving its neoclassical structure. Simulations of the Cyclone Base Case and an ASDEX Upgrade I-mode discharge show that the method reduces particle noise and substantially suppresses initialization-driven transients.

physics.plasm-ph↗

New Snake-in-the-Box Records via Snakepit Surgery and Learned Construction

The snake-in-the-box problem asks for a longest induced path in the hypercube graph $Q_n$. We find a length-191 snake in dimension $n=9$, the lowest dimension where the maximum is unknown, improving the previous record of 190 that had stood for 14 years. We also establish new lower bounds in dimensions 10-13. To find these records, we introduce snakepits, collections of disjoint snakes, to expand the search space and open new routes between snakes. This motivates our new Snakepit-in-the-Box benchmark, which seeks maximal edge counts when allowing multiple components. Finally, we introduce Beam Anchor, a search-supervised learned constructor algorithm that finds 100 inequivalent length-190 snakes in dimension 9.

cs.DM↗