Search arXiv⌕ Search

arXiv · 2610.09813

Drift Estimation for a Multi-Dimensional Lévy-Driven Stochastic Differential Equation Using Deep Neural Networks

Abstract

We develop a non-asymptotic theory for nonparametric drift estimation in discretely observed multi-dimensional Lévy-driven stochastic differential equations using sparse ReLU neural networks. Under exponential \(β\)-mixing and an exponential-tail condition on the Lévy measure, we establish an oracle inequality for the least-squares estimator. The jump component changes both the martingale structure and the concentration regime relative to diffusion models: continuous martingales are replaced by discontinuous martingales, while the relevant empirical-process fluctuations are sub-exponential rather than sub-Gaussian. We handle these difficulties without truncating the observed increments, using an exponential-supermartingale argument for compensated Poisson integrals and \(ψ_1\)-chaining. Despite the weaker concentration, the resulting statistical bound is, up to constants, of the same order as the corresponding diffusion bound. For drift functions with hierarchical compositional structure, this yields intrinsic-dimensional convergence rates. The framework allows infinite jump activity and, in some cases, infinite variation. Finally, we establish a minimax lower bound on a fixed non-trivial compound-Poisson submodel. Together with the upper bound, this shows that the estimator is minimax optimal up to logarithmic factors.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Taisei Noguchi, Teppei Ogihara. 2026-10-07. Drift Estimation for a Multi-Dimensional Lévy-Driven Stochastic Differential Equation Using Deep Neural Networks. https://arxiv.org/abs/2610.09813

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Handling Covariate Mismatch in Collaborative Linear Prediction

Training predictive models across multiple centers typically assumes that all centers collect the same set of covariates. In practice, however, they may record different features of their observations, a setting we refer to as covariate mismatch. We study linear prediction under this challenging setting, assuming center-wise MCAR missingness patterns, and develop estimators that exploit information across centers despite heterogeneous feature sets. In the low-dimensional regime, we propose a plug-in estimator of the oracle linear predictor based on component-wise aggregation of covariance and cross-moment estimates. In higher dimensions, we study an impute-then-regress strategy that first completes the missing covariates using an exchangeability-preserving imputation procedure and then fits a ridge-regularized linear model. All proposed estimators are compatible with federated learning constraints: individual-level data remain local to each center, and only aggregated quantities are exchanged. We provide asymptotic and finite-sample learning rates for our predictors, explicitly characterizing their behaviour with the global dimension, the center-specific feature partition, and the distribution of samples across centers, and validate our approach through numerical experiments.

math.ST↗

Lambda-quantiles under the microscope

We study Lambda-quantiles, a generalisation of classical quantiles in which the constant probability level $λ\in [0,1]$ is replaced by a functional parameter $Λ\colon \mathbb{R} \to [0,1]$. We consider the general case of non-monotone $Λ$, which arises naturally if closure properties of the class of corresponding Lambda-quantiles with respect to inf-aggregation or with respect to mixtures are required. As preliminary results, we characterise finiteness, constancy, and what we call the attainment property known from classical quantiles. We then consider the problem of reconstructing $Λ$ from the values of $Λ$-quantiles on a suitable family of simple distributions, showing its identifiability under mild assumptions. Next, we substantially refine several results obtained in the literature on weak upper and lower semicontinuity and on the property of convexity of the level sets with respect to mixtures, obtaining in both cases almost complete characterisations without any monotonicity assumption. We then move to the case in which $Λ$ has bounded variation, which enables us to prove a mixture representation result: any such $Λ$-quantile can be rewritten as a Lambda-quantile with an increasing functional parameter, evaluated at a mixture of the original distribution with a fixed reference distribution at a fixed weight, thus reducing the complexity of the parameter from bounded variation to monotone. Finally, we introduce and study the notion of the ordinal covariance group of a risk measure, showing that in the case of a $Λ$-quantile it coincides with the compositional invariance group of $Λ$ and with a certain group of measure-preserving transformations of the signed measure associated with $Λ$.

math.ST↗

Shape without scale: an identifiability dichotomy for a bounded tail observed through a non-additive measurement kernel

A latent severity has a bounded lower tail with density of shape alpha and scale L. It is observed only through a fixed Markov kernel K that is biased and non-additive. The relative conditional spread of K diverges at the endpoint. Our sample is i.i.d. from the marginal Q alone, with no anchoring covariate or instrument. We prove a dichotomy. The shape index alpha is identifiable: for every admissible choice of the class constants, any two observationally equivalent members of a lean class share alpha, determined by a near-endpoint expansion of Q. The rate, namely L and the fixed-scale exceedance p_tau, does not survive. There exist admissible shared class constants and two members of a smaller regularity class whose observed laws coincide exactly. Across the pair alpha agrees, whereas L and p_tau move. A degenerate Le Cam two-point bound excludes any uniformly consistent estimator of either, and pointwise consistency fails at one member. Only the rate needs an anchor. We conjecture that a known kernel family with known edge map identifies the rate fiber by fiber if and only if the family satisfies a fixed-scale injectivity clause, and we prove the sufficiency direction. In surrogate safety, uncalibrated conflict data give the shape of near-crash risk, not its absolute rate.

math.ST↗