Searcharxiv⌕ Search

arXiv subjects

Cathrine Aeckerle-Willems

Publications and source records attributed to Cathrine Aeckerle-Willems.

4 recordsLinked to original sources

Error Bounds for Importance Sampling with Estimated Proposal Distributions

Importance sampling with data-driven proposal distributions is widely used in practice. A common workflow first generates an auxiliary sample of size $N$ from an approximation of the target distribution, constructs a density estimate $\hat q$ such as a kernel density estimator (KDE), and then draws $n$ importance samples from this learned proposal. Despite its practical relevance, the theoretical properties of this hierarchical procedure remain poorly understood, since classical importance sampling theory assumes a fixed proposal. We address this gap by deriving non-asymptotic error bounds for standard, defensive, and self-normalized importance sampling estimators with random proposals. Our results separate the Monte Carlo error, scaling as $n^{-1/2}$, from the proposal approximation error measured through the mean integrated absolute and squared errors (MIAE and MISE) of $\hat q$. To obtain explicit convergence rates in $(N,n)$, we establish MIAE and MISE bounds for KDEs constructed from geometrically ergodic Markov chains in stationary and non-stationary regimes. Combining these results yields quantitative guarantees for importance sampling with KDE-based proposals. Our theory provides practical guidance for selecting defensive mixture weights in a nonparametric importance sampling framework.

math.ST↗

Concentration analysis of multivariate elliptic diffusion processes

We prove concentration inequalities and associated PAC bounds for continuous- and discrete-time additive functionals for possibly unbounded functions of multivariate, nonreversible diffusion processes. Our analysis relies on an approach via the Poisson equation allowing us to consider a very broad class of subexponentially ergodic processes. These results add to existing concentration inequalities for additive functionals of diffusion processes which have so far been only available for either bounded functions or for unbounded functions of processes from a significantly smaller class. We demonstrate the power of these exponential inequalities by two examples of very different areas. Considering a possibly high-dimensional parametric nonlinear drift model under sparsity constraints, we apply the continuous-time concentration results to validate the restricted eigenvalue condition for Lasso estimation, which is fundamental for the derivation of oracle inequalities. The results for discrete additive functionals are used to investigate the unadjusted Langevin MCMC algorithm for sampling of moderately heavy-tailed densities $π$. In particular, we provide PAC bounds for the sample Monte Carlo estimator of integrals $π(f)$ for polynomially growing functions $f$ that quantify sufficient sample and step sizes for approximation within a prescribed margin with high probability.

math.PR↗

Concentration of scalar ergodic diffusions and some statistical implications

We derive uniform concentration inequalities for continuous-time analogues of empirical processes and related stochastic integrals of scalar ergodic diffusion processes. Thereby, we lay the foundation typically required for the study of sup-norm properties of estimation procedures for a large class of diffusion processes. In the classical i.i.d. context, a key device for the statistical sup-norm analysis is provided by Talagrand-type concentration inequalities. Aiming for a parallel substitute in the diffusion framework, we present a systematic, self-contained approach to such uniform concentration inequalities via martingale approximation and moment bounds obtained by the generic chaining method. The developed machinery is of independent probabilistic interest and can serve as a starting point for investigations of other processes such as more general Markov processes, in particular multivariate or discretely observed diffusions. As a first concrete statistical application, we analyse the sup-norm error of estimating the invariant density of an ergodic diffusion via the natural local time estimator and the classical nonparametric kernel density estimator, respectively.

math.ST↗

Sup-norm adaptive simultaneous drift estimation for ergodic diffusions

We consider the question of estimating the drift and the invariant density for a large class of scalar ergodic diffusion processes, based on continuous observations, in $\sup$-norm loss. The unknown drift $b$ is supposed to belong to a nonparametric class of smooth functions of unknown order. We suggest an adaptive approach which allows to construct drift estimators attaining minimax optimal $\sup$-norm rates of convergence. In addition, we prove a Donsker theorem for the classical kernel estimator of the invariant density and establish its semiparametric efficiency. Finally, we combine both results and propose a fully data-driven bandwidth selection procedure which simultaneously yields both a rate-optimal drift estimator and an asymptotically efficient estimator of the invariant density of the diffusion. Crucial tool for our investigation are uniform exponential inequalities for empirical processes of diffusions.

math.ST↗