SearcharxivSearch

arXiv subjects

Claire Lacour

Publications and source records attributed to Claire Lacour.

At least 19 recordsLinked to original sources

Gold standard process Markovian poisoning: a semiparametric approach

We consider in this paper a stochastic process that mixes in time, according to a nonobserved stationary Markov selection process, two separate sources of randomness: i) a stationary process which distribution is accessible (gold standard); ii) a pure i.i.d. sequence which distribution is unknown (poisoning process). In this framework we propose to estimate, with two different approaches, the transition of the hidden Markov selection process along with the distribution, not supposed to belong to any parametric family, of the unknown i.i.d. sequence, under minimal (identifiability, stationarity and dependence in time) conditions. We show that both estimators provide consistent estimations of the Euclidean transition parameter, and also prove that one of them, which is $\sqrt$ n-consistent, allows to establish a functional central limit theorem about the unknown poisoning sequence cumulative distribution function. The numerical performances of our estimators are illustrated through various challenging examples.

math.ST

Nonparametric optimal density estimation for censored circular data

We consider the problem of estimating the probability density function of a circular random variable observed under censoring. To this end, we introduce a projection estimator constructed via a regression approach on linear sieves. We first establish a lower bound for the mean integrated squared error in the case of Sobolev densities, thereby identifying the minimax rate of convergence for this estimation problem. We then derive a matching upper bound for the same risk, showing that the proposed estimator attains the minimax rate when the underlying density belongs to a Sobolev class. Finally, we develop a data-driven version of the procedure that preserves this optimal rate, thus yielding an adaptive estimator. The practical performance of the method is demonstrated through simulation studies.

math.ST

Is model selection possible for the $\ell_p$-loss? PCO estimation for regression models

This paper addresses the problem of model selection in the sequence model $Y=θ+\varepsilonξ$, when $ξ$ is sub-Gaussian, for non-euclidian loss-functions. In this model, the Penalized Comparison to Overfitting procedure is studied for the weighted $\ell_p$-loss, $p\geq 1.$ Several oracle inequalities are derived from concentration inequalities for sub-Weibull variables. Using judicious collections of models and penalty terms, minimax rates of convergence are stated for Besov bodies $\mathcal{B}_{r,\infty}^s$. These results are applied to the functional model of nonparametric regression.

math.ST

Uniform Deconvolution for Poisson Point Processes

We focus on the estimation of the intensity of a Poisson process in the presence of a uniform noise. We propose a kernel-based procedure fully calibrated in theory and practice. We show that our adaptive estimator is optimal from the oracle and minimax points of view, and provide new lower bounds when the intensity belongs to a Sobolev ball. By developing the Goldenshluger-Lepski methodology in the case of deconvolution for Poisson processes, we propose an optimal data-driven selection of the kernel bandwidth. Our method is illustrated on the spatial distribution of replication origins and sequence motifs along the human genome.

stat.ME

Three rates of convergence or separation via U-statistics in a dependent framework

Despite the ubiquity of U-statistics in modern Probability and Statistics, their non-asymptotic analysis in a dependent framework may have been overlooked. In a recent work, a new concentration inequality for U-statistics of order two for uniformly ergodic Markov chains has been proved. In this paper, we put this theoretical breakthrough into action by pushing further the current state of knowledge in three different active fields of research. First, we establish a new exponential inequality for the estimation of spectra of trace class integral operators with MCMC methods. The novelty is that this result holds for kernels with positive and negative eigenvalues, which is new as far as we know. In addition, we investigate generalization performance of online algorithms working with pairwise loss functions and Markov chain samples. We provide an online-to-batch conversion result by showing how we can extract a low risk hypothesis from the sequence of hypotheses generated by any online learner. We finally give a non-asymptotic analysis of a goodness-of-fit test on the density of the invariant measure of a Markov chain. We identify some classes of alternatives over which our test based on the $L_2$ distance has a prescribed power.

math.ST

Semiparametric inference for mixtures of circular data

We consider X 1 ,. .. , X n a sample of data on the circle S 1 , whose distribution is a twocomponent mixture. Denoting R and Q two rotations on S 1 , the density of the X i 's is assumed to be g(x) = pf (R --1 x) + (1 -- p)f (Q --1 x), where p $\in$ (0, 1) and f is an unknown density on the circle. In this paper we estimate both the parametric part $θ$ = (p, R, Q) and the nonparametric part f. The specific problems of identifiability on the circle are studied. A consistent estimator of $θ$ is introduced and its asymptotic normality is proved. We propose a Fourier-based estimator of f with a penalized criterion to choose the resolution level. We show that our adaptive estimator is optimal from the oracle and minimax points of view when the density belongs to a Sobolev ball. Our method is illustrated by numerical simulations.

math.ST

Concentration inequality for U-statistics of order two for uniformly ergodic Markov chains

We prove a new concentration inequality for U-statistics of order two for uniformly ergodic Markov chains. Working with bounded and $π$-canonical kernels, we show that we can recover the convergence rate of Arcones and Gin{é} who proved a concentration result for U-statistics of independent random variables and canonical kernels. Our result allows for a dependence of the kernels $h_{i,j}$ with the indexes in the sums, which prevents the use of standard blocking tools. Our proof relies on an inductive analysis where we use martingale techniques, uniform ergodicity, Nummelin splitting and Bernstein's type inequality. Assuming further that the Markov chain starts from its invariant distribution, we prove a Bernstein-type concentration inequality that provides sharper convergence rate for small variance terms.

math.PR

Adaptive greedy algorithm for moderately large dimensions in kernel conditional density estimation

This paper studies the estimation of the conditional density f (x, $\times$) of Y i given X i = x, from the observation of an i.i.d. sample (X i , Y i) $\in$ R d , i = 1,. .. , n. We assume that f depends only on r unknown components with typically r d. We provide an adaptive fully-nonparametric strategy based on kernel rules to estimate f. To select the bandwidth of our kernel rule, we propose a new fast iterative algorithm inspired by the Rodeo algorithm (Wasserman and Lafferty (2006)) to detect the sparsity structure of f. More precisely, in the minimax setting, our pointwise estimator, which is adaptive to both the regularity and the sparsity, achieves the quasi-optimal rate of convergence. Its computational complexity is only O(dn log n).

math.ST

Adaptive Estimation of Nonparametric Geometric Graphs

This article studies the recovery of graphons when they are convolution kernels on compact (symmetric) metric spaces. This case is of particular interest since it covers the situation where the probability of an edge depends only on some unknown nonparametric function of the distance between latent points, referred to as Nonparametric Geometric Graphs (NGG). In this setting, adaptive estimation of NGG is possible using a spectral procedure combined with a Goldenshluger-Lepski adaptation method. The latent spaces covered by our framework encompass (among others) compact symmetric spaces of rank one, namely real spheres and projective spaces. For these latter, explicit computations of the eigen-basis and of the model complexity can be achieved, leading to quantitative non-asymptotic results. The time complexity of our method scales cubicly in the size of the graph and exponentially in the regularity of the graphon. Hence, this paper offers an algorithmically and theoretically efficient procedure to estimate smooth NGG. As a by product, this paper shows a non-asymptotic concentration result on the spectrum of integral operators defined by symmetric kernels (not necessarily positive).

math.ST

Numerical performance of Penalized Comparison to Overfitting for multivariate kernel density estimation

Kernel density estimation is a well known method involving a smoothing parameter (the bandwidth) that needs to be tuned by the user. Although this method has been widely used the bandwidth selection remains a challenging issue in terms of balancing algorithmic performance and statistical relevance. The purpose of this paper is to compare a recently developped bandwidth selection method for kernel density estimation to those which are commonly used by now (at least those which are implemented in the R-package). This new method is called Penalized Comparison to Overfitting (PCO). It has been proposed by some of the authors of this paper in a previous work devoted to its statistical relevance from a purely theoretical perspective. It is compared here to other usual bandwidth selection methods for univariate and also multivariate kernel density estimation on the basis of intensive simulation studies. In particular, cross-validation and plug-in criteria are numerically investigated and compared to PCO. The take home message is that PCO can outperform the classical methods without algorithmic additionnal cost.

math.ST

Estimator selection: a new method with applications to kernel density estimation

Estimator selection has become a crucial issue in non parametric estimation. Two widely used methods are penalized empirical risk minimization (such as penalized log-likelihood estimation) or pairwise comparison (such as Lepski's method). Our aim in this paper is twofold. First we explain some general ideas about the calibration issue of estimator selection methods. We review some known results, putting the emphasis on the concept of minimal penalty which is helpful to design data-driven selection criteria. Secondly we present a new method for bandwidth selection within the framework of kernel density density estimation which is in some sense intermediate between these two main methods mentioned above. We provide some theoretical results which lead to some fully data-driven selection strategy.

math.ST

Minimal penalty for Goldenshluger-Lepski method

This paper is concerned with adaptive nonparametric estimation using the Goldenshluger-Lepski selection method. This estimator selection method is based on pairwise comparisons between estimators with respect to some loss function. The method also involves a penalty term that typically needs to be large enough in order that the method works (in the sense that one can prove some oracle type inequality for the selected estimator). In the case of density estimation with kernel estimators and a quadratic loss, we show that the procedure fails if the penalty term is chosen smaller than some critical value for the penalty: the minimal penalty. More precisely we show that the quadratic risk of the selected estimator explodes when the penalty is below this critical value while it stays under control when the penalty is above this critical value. This kind of phase transition phenomenon for penalty calibration has already been observed and proved for penalized model selection methods in various contexts but appears here for the first time for the Goldenshluger-Lepski pairwise comparison method. Some simulations illustrate the theoretical results and lead to some hints on how to use the theory to calibrate the method in practice.

math.ST

Minimax adaptive estimation of nonparametric hidden Markov models

We consider stationary hidden Markov models with finite state space and nonparametric modeling of the emission distributions. It has remained unknown until very recently that such models are identifiable. In this paper, we propose a new penalized least-squares esti-mator for the emission distributions which is statistically optimal and practically tractable. We prove a non asymptotic oracle inequality for our nonparametric estimator of the emission distributions. A consequence is that this new estimator is rate minimax adaptive up to a logarithmic term. Our methodology is based on projections of the emission distributions onto nested subspaces of increasing complexity. The popular spectral estimators are unable to achieve the optimal rate but may be used as initial points in our procedure. Simulations are given that show the improvement obtained when applying the least-squares minimization consecutively to the spectral estimation.

math.ST

Adaptive pointwise estimation of conditional density function

In this paper we consider the problem of estimating $f$, the conditional density of $Y$ given $X$, by using an independent sample distributed as $(X,Y)$ in the multivariate setting. We consider the estimation of $f(x,.)$ where $x$ is a fixed point. We define two different procedures of estimation, the first one using kernel rules, the second one inspired from projection methods. Both adapted estimators are tuned by using the Goldenshluger and Lepski methodology. After deriving lower bounds, we show that these procedures satisfy oracle inequalities and are optimal from the minimax point of view on anisotropic H{ö}lder balls. Furthermore, our results allow us to measure precisely the influence of $\mathrm{f}\_X(x)$ on rates of convergence, where $\mathrm{f}\_X$ is the density of $X$. Finally, some simulations illustrate the good behavior of our tuned estimates in practice.

math.ST

Goodness-of-fit test for noisy directional data

We consider spherical data $X_i$ noised by a random rotation $\varepsilon_i\in$ SO(3) so that only the sample $Z_i=\varepsilon_iX_i$, $i=1,\dots, N$ is observed. We define a nonparametric test procedure to distinguish $H_0:$ ''the density $f$ of $X_i$ is the uniform density $f_0$ on the sphere'' and $H_1:$ ''$\|f-f_0\|_2^2\geq \Cψ_N$ and $f$ is in a Sobolev space with smoothness $s$''. For a noise density $f_\varepsilon$ with smoothness index $ν$, we show that an adaptive procedure (i.e. $s$ is not assumed to be known) cannot have a faster rate of separation than $ψ_N^{ad}(s)=(N/\sqrt{\log\log(N)})^{-2s/(2s+2ν+1)}$ and we provide a procedure which reaches this rate. We also deal with the case of super smooth noise. We illustrate the theory by implementing our test procedure for various kinds of noise on SO(3) and by comparing it to other procedures. Applications to real data in astrophysics and paleomagnetism are provided.

math.ST

Adaptive pointwise estimation for pure jump Lévy processes

This paper is concerned with adaptive kernel estimation of the Lévy density N(x) for bounded-variation pure-jump Lévy processes. The sample path is observed at n discrete instants in the "high frequency" context (Δ= Δ(n) tends to zero while nΔtends to infinity). We construct a collection of kernel estimators of the function g(x)=xN(x) and propose a method of local adaptive selection of the bandwidth. We provide an oracle inequality and a rate of convergence for the quadratic pointwise risk. This rate is proved to be the optimal minimax rate. We give examples and simulation results for processes fitting in our framework. We also consider the case of irregular sampling.

math.ST

Least squares type estimation of the transition density of a particular hidden Markov chain

In this paper, we study the following model of hidden Markov chain: $Y_i=X_i+ε_i$, $i=1,...,n+1$ with $(X_i)$ a real-valued stationary Markov chain and $(ε_i)_{1\leq i\leq n+1}$ a noise having a known distribution and independent of the sequence $(X_i)$. We present an estimator of the transition density obtained by minimization of an original contrast that takes advantage of the regressive aspect of the problem. It is selected among a collection of projection estimators with a model selection method. The $L^2$-risk and its rate of convergence are evaluated for ordinary smooth noise and some simulations illustrate the method. We obtain uniform risk bounds over classes of Besov balls. In addition our estimation procedure requires no prior knowledge of the regularity of the true transition. Finally, our estimator permits to avoid the drawbacks of quotient estimators.

math.ST

Nonparametric estimation of the stationary density and the transition density of a Markov chain

In this paper, we study first the problem of nonparametric estimation of the stationary density $f$ of a discrete-time Markov chain $(X_i)$. We consider a collection of projection estimators on finite dimensional linear spaces. We select an estimator among the collection by minimizing a penalized contrast. The same technique enables to estimate the density $g$ of $(X_i, X_{i+1})$ and so to provide an adaptive estimator of the transition density $π=g/f$. We give bounds in $L^2$ norm for these estimators and we show that they are adaptive in the minimax sense over a large class of Besov spaces. Some examples and simulations are also provided.

math.ST