SearcharxivSearch

arXiv subjects

Nicolas Klutchnikoff

Publications and source records attributed to Nicolas Klutchnikoff.

17 recordsLinked to original sources

On the Dirichlet-kernel Gasser--Müller estimator and its competitors for fixed design regression on the simplex

A Dirichlet-kernel Gasser-Müller (D-GM) estimator is introduced for fixed design regression on the simplex, extending the univariate analog due to Chen [Statist. Sinica, vol. 10(1) (2000), pp. 73-91]. Its pointwise bias and variance, asymptotic normality, and mean integrated squared error are investigated. Some simulation experiments are conducted to compare its small-sample performance with that of two recently proposed alternatives: the Dirichlet-kernel Nadaraya-Watson (D-NW) and local linear (D-LL) estimators. The simulation results reveal that the D-LL estimator is best among the D-LL, D-NW, and D-GM estimators and that the proposed D-GM estimator is worst. A real data analysis is also reported for the GEMAS dataset to analyze the relationship between soil composition and pH levels across various agricultural and grazing lands in Europe.

math.ST

Adaptive functional principal components analysis

Functional data analysis almost always involves smoothing discrete observations into curves, because they are never observed in continuous time and rarely without error. Although smoothing parameters affect the subsequent inference, data-driven methods for selecting these parameters are not well-developed, frustrated by the difficulty of using all the information shared by curves while being computationally efficient. On the one hand, smoothing individual curves in an isolated, albeit sophisticated way, ignores useful signals present in other curves. On the other hand, bandwidth selection by automatic procedures such as cross-validation after pooling all the curves together quickly become computationally unfeasible due to the large number of data points. In this paper we propose a new data-driven, adaptive kernel smoothing, specifically tailored for functional principal components analysis through the derivation of sharp, explicit risk bounds for the eigen-elements. The minimization of these quadratic risk bounds provide refined, yet computationally efficient bandwidth rules for each eigen-element separately. Both common and independent design cases are allowed. Rates of convergence for the estimators are derived. An extensive simulation study, designed in a versatile manner to closely mimic the characteristics of real data sets supports our methodological contribution. An illustration on a real data application is provided.

stat.ME

A new adaptive local polynomial density estimation procedure on complicated domains

This paper presents a novel approach for pointwise estimation of multivariate density functions on known domains of arbitrary dimensions using nonparametric local polynomial estimators. Our method is highly flexible, as it applies to both simple domains, such as open connected sets, and more complicated domains that are not star-shaped around the point of estimation. This enables us to handle domains with sharp concavities, holes, and local pinches, such as polynomial sectors. Additionally, we introduce a data-driven selection rule based on the general ideas of Goldenshluger and Lepski. Our results demonstrate that the local polynomial estimators are minimax under a $L^2$ risk across a wide range of Hölder-type functional classes. In the adaptive case, we provide oracle inequalities and explicitly determine the convergence rate of our statistical procedure. Simulations on polynomial sectors show that our oracle estimates outperform those of the most popular alternative method, found in the sparr package for the R software. Our statistical procedure is implemented in an online R package which is readily accessible.

math.ST

Adaptive estimation of irregular mean and covariance functions

Nonparametric estimators for the mean and the covariance functions of functional data are proposed. The setup covers a wide range of practical situations. The random trajectories are, not necessarily differentiable, have unknown regularity, and are measured with error at discrete design points. The measurement error could be heteroscedastic. The design points could be either randomly drawn or common for all curves. The estimators depend on the local regularity of the stochastic process generating the functional data. We consider a simple estimator of this local regularity which exploits the replication and regularization features of functional data. Next, we use the ``smoothing first, then estimate'' approach for the mean and the covariance functions. They can be applied with both sparsely or densely sampled curves, are easy to calculate and to update, and perform well in simulations. Simulations built upon an example of real data set, illustrate the effectiveness of the new approach.

math.ST

Optimal 1-Wasserstein Distance for WGANs

The mathematical forces at work behind Generative Adversarial Networks raise challenging theoretical issues. Motivated by the important question of characterizing the geometrical properties of the generated distributions, we provide a thorough analysis of Wasserstein GANs (WGANs) in both the finite sample and asymptotic regimes. We study the specific case where the latent space is univariate and derive results valid regardless of the dimension of the output space. We show in particular that for a fixed sample size, the optimal WGANs are closely linked with connected paths minimizing the sum of the squared Euclidean distances between the sample points. We also highlight the fact that WGANs are able to approach (for the 1-Wasserstein distance) the target distribution as the sample size tends to infinity, at a given convergence rate and provided the family of generative Lipschitz functions grows appropriately. We derive in passing new results on optimal transport theory in the semi-discrete setting.

stat.ML

Learning the regularity of multivariate functional data

Combining information both within and between sample realizations, we propose a simple estimator for the local regularity of surfaces in the functional data framework. The independently generated surfaces are measured with errors at possibly random discrete times. Non-asymptotic exponential bounds for the concentration of the regularity estimators are derived. An indicator for anisotropy is proposed and an exponential bound of its risk is derived. Two applications are proposed. We first consider the class of multi-fractional, bi-dimensional, Brownian sheets with domain deformation, and study the nonparametric estimation of the deformation. As a second application, we build minimax optimal, bivariate kernel estimators for the reconstruction of the surfaces.

math.ST

Minimax properties of Dirichlet kernel density estimators

This paper considers the asymptotic behavior in $β$-Hölder spaces, and under $L^p$ losses, of a Dirichlet kernel density estimator proposed by Aitchison and Lauder (1985) for the analysis of compositional data. In recent work, Ouimet and Tolosana-Delgado (2022) established the uniform strong consistency and asymptotic normality of this estimator. As a complement, it is shown here that the Aitchison-Lauder estimator can achieve the minimax rate asymptotically for a suitable choice of bandwidth whenever $(p,β) \in [1, 3) \times (0, 2]$ or $(p, β) \in \mathcal{A}_d$, where $\mathcal{A}_d$ is a specific subset of $[3, 4) \times (0, 2]$ that depends on the dimension $d$ of the Dirichlet kernel. It is also shown that this estimator cannot be minimax when either $p \in [4, \infty)$ or $β\in (2, \infty)$. These results extend to the multivariate case, and also rectify in a minor way, earlier findings of Bertin and Klutchnikoff (2011) concerning the minimax properties of Beta kernel estimators.

math.ST

Statistical analysis of a hierarchical clustering algorithm with outliers

It is well known that the classical single linkage algorithm usually fails to identify clusters in the presence of outliers. In this paper, we propose a new version of this algorithm, and we study its mathematical performances. In particular, we establish an oracle type inequality which ensures that our procedure allows to recover the clusters with large probability under minimal assumptions on the distribution of the outliers. We deduce from this inequality the consistency and some rates of convergence of our algorithm for various situations. Performances of our approach is also assessed through simulation studies and a comparison with classical clustering algorithms on simulated data is also presented.

math.ST

Learning the smoothness of noisy curves with application to online curve estimation

Combining information both within and across trajectories, we propose a simple estimator for the local regularity of the trajectories of a stochastic process. Independent trajectories are measured with errors at randomly sampled time points. Non-asymptotic bounds for the concentration of the estimator are derived. Given the estimate of the local regularity, we build a nearly optimal local polynomial smoother from the curves from a new, possibly very large sample of noisy trajectories. We derive non-asymptotic pointwise risk bounds uniformly over the new set of curves. Our estimates perform well in simulations. Real data sets illustrate the effectiveness of the new approaches.

math.ST

Clustering multivariate functional data using unsupervised binary trees

We propose a model-based clustering algorithm for a general class of functional data for which the components could be curves or images. The random functional data realizations could be measured with error at discrete, and possibly random, points in the definition domain. The idea is to build a set of binary trees by recursive splitting of the observations. The number of groups are determined in a data-driven way. The new algorithm provides easily interpretable results and fast predictions for online data sets. Results on simulated datasets reveal good performance in various complex settings. The methodology is applied to the analysis of vehicle trajectories on a German roundabout.

stat.ML

Adaptive regression with Brownian path covariate

This paper deals with estimation with functional covariates. More precisely, we aim at estimating the regression function $m$ of a continuous outcome $Y$ against a standard Wiener coprocess $W$. Following Cadre and Truquet (2015) and Cadre, Klutchnikoff, and Massiot (2017) the Wiener-Itô decomposition of $m(W)$ is used to construct a family of estimators. The minimax rate of convergence over specific smoothness classes is obtained. A data-driven selection procedure is defined following the ideas developed by Goldenshluger and Lepski (2011). An oracle-type inequality is obtained which leads to adaptive results.

math.ST

Adaptive Density Estimation on Bounded Domains

We study the estimation, in Lp-norm, of density functions defined on [0,1]^d. We construct a new family of kernel density estimators that do not suffer from the so-called boundary bias problem and we propose a data-driven procedure based on the Goldenshluger and Lepski approach that jointly selects a kernel and a bandwidth. We derive two estimators that satisfy oracle-type inequalities. They are also proved to be adaptive over a scale of anisotropic or isotropic Sobolev-Slobodetskii classes (which are particular cases of Besov or Sobolev classical classes). The main interest of the isotropic procedure is to obtain adaptive results without any restriction on the smoothness parameter.

math.ST

Kernel estimation of the intensity of Cox processes

Counting processes often written $N=(N_t)_{t\in\mathbb{R}^+}$ are used in several applications of biostatistics, notably for the study of chronic diseases. In the case of respiratory illness it is natural to suppose that the count of the visits of a patient can be described by such a process which intensity depends on environmental covariates. Cox processes (also called doubly stochastic Poisson processes) allows to model such situations. The random intensity then writes $λ(t)=θ(t,Z_t)$ where $θ$ is a non-random function, $t\in\mathbb{R}^+$ is the time variable and $(Z_t)_{t\in\mathbb{R}^+}$ is the $d$-dimensional covariates process. For a longitudinal study over $n$ patients, we observe $(N_t^k,Z_t^k)_{t\in\mathbb{R}^+}$ for $k=1,\ldots,n$. The intention is to estimate the intensity of the process using these observations and to study the properties of this estimator.

math.ST

Pointwise Adaptive Estimation of the MarginalDensity of a Weakly Dependent Process

This paper is devoted to the estimation of the common marginal density function of weakly dependent processes. The accuracy of estimation is measured using pointwise risks. We propose a datadriven procedure using kernel rules. The bandwidth is selected using the approach of Goldenshluger and Lepski and we prove that the resulting estimator satisfies an oracle type inequality. The procedure is also proved to be adaptive (in a minimax framework) over a scale of Hölder balls for several types of dependence: stong mixing processes, $λ$-dependent processes or i.i.d. sequences can be considered using a single procedure of estimation. Some simulations illustrate the performance of the proposed method.

math.ST

On clustering procedures and nonparametric mixture estimation

This paper deals with nonparametric estimation of conditional den-sities in mixture models in the case when additional covariates are available. The proposed approach consists of performing a prelim-inary clustering algorithm on the additional covariates to guess the mixture component of each observation. Conditional densities of the mixture model are then estimated using kernel density estimates ap-plied separately to each cluster. We investigate the expected L 1 -error of the resulting estimates and derive optimal rates of convergence over classical nonparametric density classes provided the clustering method is accurate. Performances of clustering algorithms are measured by the maximal misclassification error. We obtain upper bounds of this quantity for a single linkage hierarchical clustering algorithm. Lastly, applications of the proposed method to mixture models involving elec-tricity distribution data and simulated data are presented.

math.ST

Minimax properties of beta kernel density estimators

In this paper, we are interested in the study of beta kernel estimators from an asymptotic minimax point of view. It is well known that beta kernel estimators are, on the contrary of classical kernel estimators, "free of boundary effect" and thus are very useful in practice. The goal of this paper is to prove that there is a price to pay: for very regular functions or for certain losses, these estimators are not minimax. Nevertheless they are minimax for classical regularities such as regularity of order two or less than two, supposed commonly in the practice and for some classical losses.

math.ST