Searcharxiv⌕ Search

arXiv · 2610.09813

Drift Estimation for a Multi-Dimensional Lévy-Driven Stochastic Differential Equation Using Deep Neural Networks

Abstract

We develop a non-asymptotic theory for nonparametric drift estimation in discretely observed multi-dimensional Lévy-driven stochastic differential equations using sparse ReLU neural networks. Under exponential \(β\)-mixing and an exponential-tail condition on the Lévy measure, we establish an oracle inequality for the least-squares estimator. The jump component changes both the martingale structure and the concentration regime relative to diffusion models: continuous martingales are replaced by discontinuous martingales, while the relevant empirical-process fluctuations are sub-exponential rather than sub-Gaussian. We handle these difficulties without truncating the observed increments, using an exponential-supermartingale argument for compensated Poisson integrals and \(ψ_1\)-chaining. Despite the weaker concentration, the resulting statistical bound is, up to constants, of the same order as the corresponding diffusion bound. For drift functions with hierarchical compositional structure, this yields intrinsic-dimensional convergence rates. The framework allows infinite jump activity and, in some cases, infinite variation. Finally, we establish a minimax lower bound on a fixed non-trivial compound-Poisson submodel. Together with the upper bound, this shows that the estimator is minimax optimal up to logarithmic factors.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Taisei Noguchi, Teppei Ogihara. 2026-10-07. Drift Estimation for a Multi-Dimensional Lévy-Driven Stochastic Differential Equation Using Deep Neural Networks. https://arxiv.org/abs/2610.09813

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Lambda-quantiles under the microscope

We study Lambda-quantiles, a generalisation of classical quantiles in which the constant probability level $λ\in [0,1]$ is replaced by a functional parameter $Λ\colon \mathbb{R} \to [0,1]$. We consider the general case of non-monotone $Λ$, which arises naturally if closure properties of the class of corresponding Lambda-quantiles with respect to inf-aggregation or with respect to mixtures are required. As preliminary results, we characterise finiteness, constancy, and what we call the attainment property known from classical quantiles. We then consider the problem of reconstructing $Λ$ from the values of $Λ$-quantiles on a suitable family of simple distributions, showing its identifiability under mild assumptions. Next, we substantially refine several results obtained in the literature on weak upper and lower semicontinuity and on the property of convexity of the level sets with respect to mixtures, obtaining in both cases almost complete characterisations without any monotonicity assumption. We then move to the case in which $Λ$ has bounded variation, which enables us to prove a mixture representation result: any such $Λ$-quantile can be rewritten as a Lambda-quantile with an increasing functional parameter, evaluated at a mixture of the original distribution with a fixed reference distribution at a fixed weight, thus reducing the complexity of the parameter from bounded variation to monotone. Finally, we introduce and study the notion of the ordinal covariance group of a risk measure, showing that in the case of a $Λ$-quantile it coincides with the compositional invariance group of $Λ$ and with a certain group of measure-preserving transformations of the signed measure associated with $Λ$.

math.ST↗

Shape without scale: an identifiability dichotomy for a bounded tail observed through a non-additive measurement kernel

A latent severity has a bounded lower tail with density of shape alpha and scale L. It is observed only through a fixed Markov kernel K that is biased and non-additive. The relative conditional spread of K diverges at the endpoint. Our sample is i.i.d. from the marginal Q alone, with no anchoring covariate or instrument. We prove a dichotomy. The shape index alpha is identifiable: for every admissible choice of the class constants, any two observationally equivalent members of a lean class share alpha, determined by a near-endpoint expansion of Q. The rate, namely L and the fixed-scale exceedance p_tau, does not survive. There exist admissible shared class constants and two members of a smaller regularity class whose observed laws coincide exactly. Across the pair alpha agrees, whereas L and p_tau move. A degenerate Le Cam two-point bound excludes any uniformly consistent estimator of either, and pointwise consistency fails at one member. Only the rate needs an anchor. We conjecture that a known kernel family with known edge map identifies the rate fiber by fiber if and only if the family satisfies a fixed-scale injectivity clause, and we prove the sufficiency direction. In surrogate safety, uncalibrated conflict data give the shape of near-crash risk, not its absolute rate.

math.ST↗

Decision-Sufficient Posterior Approximation

We investigate the consequences of requiring a posterior approximation to preserve a specified downstream decision problem. A target posterior $P$ and loss determine a regret geometry on actions, a baseline approximation $Q_0$ determines the forward-Kullback-Leibler information required to induce action changes, and a restricted approximation family $\mathcal{Q}$ determines which such changes are available. Contracting KL divergence over Bayes-action fibers gives exact distances to decision adequacy and decision failure together with the least-informative posterior deformations that reach either side of the decision boundary. In regular finite-dimensional problems, the target and baseline constructions have quadratic local limits: a target regret Hessian $G$ and a baseline information metric $J_I$ . Their generalized eigenproblem $Gv = γJ_Iv$ orders local decision directions by regret consequence per unit information cost and induces a tolerance-dependent effective dimension. For restricted approximation families, the tangent image separates decision coverage from information efficiency: a family may miss consequential decision directions, or it may realize reachable directions only at excess Fisher cost. The resulting framework provides decision-relative criteria for comparing and designing posterior approximation families.

math.ST↗