SearcharxivSearch

arXiv subjects

Marie-Luce Taupin

Publications and source records attributed to Marie-Luce Taupin.

17 recordsLinked to original sources

Faster Learning under Relaxed Local Differential Privacy

We consider density estimation under the relaxed local differential privacy condition that the privatized distributions are $α$-close in total variation distance. We show that adding independent noise with a convenient symmetrized Gamma distribution to each sensitive observation attains the $α$-TV-LDP. We prove that the deconvolution estimator of $r$-Sobolev smooth functions attains the pointwise rate $(nα)^{-\frac{2r-1}{2r}}$ up to log factors which is faster than $(nα^2)^{-\frac{2r-1}{2r+1}}$ under the classical $α$-LDP and closer to the nonprivate minimax rate $n^{-\frac{2r-1}{2r}}$. Next, we use a Goldenshluger-Lepski procedure to build a free of the smoothness adaptive procedure and show optimality of our rates in the convolution model of our privatisation scheme. We illustrate the benefits of this simple privacy mechanism by implementing a neural network estimator which does not need to add more noise in the optimization steps. Numerical results show significant improvement of the estimation rate over the Laplace and the private-SGD mechanisms.

math.ST

Asymptotic confidence interval for R2 in multiple linear regression

Following White's approach of robust multiple linear regression, we give asymptotic confidence intervals for the multiple correlation coefficient R2 under minimal moment conditions. We also give the asymptotic joint distribution of the empirical estimators of the individual R2's. Through different sets of simulations, we show that the procedure is indeed robust (contrary to the procedure involving the near exact distribution of the empirical estimator of R2 is the multivariate Gaussian case) and can be also applied to count linear regression.

math.ST

RKHSMetaMod: An R package to estimate the Hoeffding decomposition of a complex model by solving RKHS ridge group sparse optimization problem

In this paper, we propose an R package, called RKHSMetaMod, that implements a procedure for estimating a meta-model of a complex model. The meta-model approximates the Hoeffding decomposition of the complex model and allows us to perform sensitivity analysis on it. It belongs to a reproducing kernel Hilbert space that is constructed as a direct sum of Hilbert spaces. The estimator of the meta-model is the solution of a penalized empirical least-squares minimization with the sum of the Hilbert norm and the empirical L^2-norm. This procedure, called RKHS ridge group sparse, allows both to select and estimate the terms in the Hoeffding decomposition, and therefore, to select and estimate the Sobol indices that are non-zero. The RKHSMetaMod package provides an interface from R statistical computing environment to the C++ libraries Eigen and GSL. In order to speed up the execution time and optimize the storage memory, except for a function that is written in R, all of the functions of this package are written using the efficient C++ libraries through RcppEigen and RcppGSL packages. These functions are then interfaced in the R environment in order to propose a user-friendly package.

stat.ML

Risk upper bounds for RKHS ridge group sparse estimator in the regression model with non-Gaussian and non-bounded error

We consider the problem of estimating a meta-model of an unknown regression model with non-Gaussian and non-bounded error. The meta-model belongs to a reproducing kernel Hilbert space constructed as a direct sum of Hilbert spaces leading to an additive decomposition including the variables and interactions between them. The estimator of this meta-model is calculated by minimizing an empirical least-squares criterion penalized by the sum of the Hilbert norm and the empirical $L^2$-norm. In this context, the upper bounds of the empirical $L^2$ risk and the $L^2$ risk of the estimator are established.

math.ST

Metamodel construction for sensitivity analysis

We propose to estimate a metamodel and the sensitivity indices of a complex model m in the Gaussian regression framework. Our approach combines methods for sensitivity analysis of complex models and statistical tools for sparse non-parametric estimation in multivariate Gaussian regression model. It rests on the construction of a metamodel for aproximating the Hoeffding-Sobol decomposition of m. This metamodel belongs to a reproducing kernel Hilbert space constructed as a direct sum of Hilbert spaces leading to a functional ANOVA decomposition. The estimation of the metamodel is carried out via a penalized least-squares minimization allowing to select the subsets of variables that contribute to predict the output. It allows to estimate the sensitivity indices of m. We establish an oracle-type inequality for the risk of the estimator, describe the procedure for estimating the metamodel and the sensitivity indices, and assess the performances of the procedure via a simulation study.

math.ST

Sensitivity analysis of spatio-temporal models describing nitrogen transfers, transformations and losses at the landscape scale

Modelling complex systems such as agroecosystems often requires the quantification of a large number of input factors. Sensitivity analyses are useful to determine the appropriate spatial and temporal resolution of models and to reduce the number of factors to be measured or estimated accurately. Comprehensive spatial and temporal sensitivity analyses were applied to the NitroScape model, a deterministic spatially distributed model describing nitrogen transfers and transformations in rural landscapes. Simulations were led on a theoretical landscape that represented five years of intensive farm management and covering an area of $3\, km^2$. Cluster analyses were applied to summarize the results of the sensitivity analysis on the ensemble of model outputs. The methodology we applied is useful to synthesize sensitivity analyses of models with multiple space-time input and output variables and could be ported to other models than NitroScape.

stat.AP

Model selection in logistic regression

This paper is devoted to model selection in logistic regression. We extend the model selection principle introduced by Birgé and Massart (2001) to logistic regression model. This selection is done by using penalized maximum likelihood criteria. We propose in this context a completely data-driven criteria based on the slope heuristics. We prove non asymptotic oracle inequalities for selected estimators. Theoretical results are illustrated through simulation studies.

math.ST

Adaptive kernel estimation of the baseline function in the Cox model, with high-dimensional covariates

The aim of this article is to propose a novel kernel estimator of the baseline function in a general high-dimensional Cox model, for which we derive non-asymptotic rates of convergence. To construct our estimator, we first estimate the regression parameter in the Cox model via a Lasso procedure. We then plug this estimator into the classical kernel estimator of the baseline function, obtained by smoothing the so-called Breslow estimator of the cumulative baseline function. We propose and study an adaptive procedure for selecting the bandwidth, in the spirit of Gold-enshluger and Lepski (2011). We state non-asymptotic oracle inequalities for the final estimator, which reveal the reduction of the rates of convergence when the dimension of the covariates grows.

stat.AP

Adaptive estimation of the baseline hazard function in the Cox model by model selection, with high-dimensional covariates

The purpose of this article is to provide an adaptive estimator of the baseline function in the Cox model with high-dimensional covariates. We consider a two-step procedure : first, we estimate the regression parameter of the Cox model via a Lasso procedure based on the partial log-likelihood, secondly, we plug this Lasso estimator into a least-squares type criterion and then perform a model selection procedure to obtain an adaptive penalized contrast estimator of the baseline function. Using non-asymptotic estimation results stated for the Lasso estimator of the regression parameter, we establish a non-asymptotic oracle inequality for this penalized contrast estimator of the baseline function, which highlights the discrepancy of the rate of convergence when the dimension of the covariates increases.

math.ST

Estimation in autoregressive model with measurement error

Consider an autoregressive model with measurement error: we observe $Z_i=X_i+ε_i$, where $X_i$ is a stationary solution of the equation $X_i=f_{θ^0}(X_{i-1})+ξ_i$. The regression function $f_{θ^0}$ is known up to a finite dimensional parameter $θ^0$. The distributions of $X_0$ and $ξ_1$ are unknown whereas the distribution of $ε_1$ is completely known. We want to estimate the parameter $θ^0$ by using the observations $Z_0,..,Z_n$. We propose an estimation procedure based on a modified least square criterion involving a weight function $w$, to be suitably chosen. We give upper bounds for the risk of the estimator, which depend on the smoothness of the errors density $f_ε$ and on the smoothness properties of $w f_θ$.

math.ST

New $M$-estimators in semi-parametric regression with errors in variables

In the regression model with errors in variables, we observe $n$ i.i.d. copies of $(Y,Z)$ satisfying $Y=f_{θ^0}(X)+ξ$ and $Z=X+ε$ involving independent and unobserved random variables $X,ξ,ε$ plus a regression function $f_{θ^0}$, known up to a finite dimensional $θ^0$. The common densities of the $X_i$'s and of the $ξ_i$'s are unknown, whereas the distribution of $ε$ is completely known. We aim at estimating the parameter $θ^0$ by using the observations $(Y_1,Z_1),...,(Y_n,Z_n)$. We propose an estimation procedure based on the least square criterion $\tilde{S}_{θ^0,g}(θ)=\m athbb{E}_{θ^0,g}[((Y-f_θ(X))^2w(X)]$ where $w$ is a weight function to be chosen. We propose an estimator and derive an upper bound for its risk that depends on the smoothness of the errors density $p_ε$ and on the smoothness properties of $w(x)f_θ(x)$. Furthermore, we give sufficient conditions that ensure that the parametric rate of convergence is achieved. We provide practical recipes for the choice of $w$ in the case of nonlinear regression functions which are smooth on pieces allowing to gain in the order of the rate of convergence, up to the parametric rate in some cases. We also consider extensions of the estimation procedure, in particular, when a choice of $w_θ$ depending on $θ$ would be more appropriate.

math.ST

Adaptive density estimation for general ARCH models

We consider a model $Y\_t=σ\_tη\_t$ in which $(σ\_t)$ is not independent of the noise process $(η\_t)$, but $σ\_t$ is independent of $η\_t$ for each $t$. We assume that $(σ\_t)$ is stationary and we propose an adaptive estimator of the density of $\ln(σ^2\_t)$ based on the observations $Y\_t$. Under various dependence structures, the rates of this nonparametric estimator coincide with the minimax rates obtained in the i.i.d. case when $(σ\_t)$ and $(η\_t)$ are independent, in all cases where these minimax rates are known. The results apply to various linear and non linear ARCH processes.

math.ST

Semi-parametric estimation of the hazard function in a model with covariate measurement error

We consider a model where the failure hazard function, conditional on a covariate $Z$ is given by $R(t,θ^0|Z)=η\_{γ^0}(t)f\_{β^0}(Z)$, with $θ^0=(β^0,γ^0)^\top\in \mathbb{R}^{m+p}$. The baseline hazard function $η\_{γ^0}$ and relative risk $f\_{β^0}$ belong both to parametric families. The covariate $Z$ is measured through the error model $U=Z+ε$ where $ε$ is independent from $Z$, with known density $f\_ε$. We observe a $n$-sample $(X\_i, D\_i, U\_i)$, $i=1,...,n$, where $X\_i$ is the minimum between the failure time and the censoring time, and $D\_i$ is the censoring indicator. We aim at estimating $θ^0$ in presence of the unknown density $g$. Our estimation procedure based on least squares criterion provide two estimators. The first one minimizes an estimation of the least squares criterion where $g$ is estimated by density deconvolution. Its rate depends on the smoothnesses of $f\_ε$ and $f\_β(z)$ as a function of $z$,. We derive sufficient conditions that ensure the $\sqrt{n}$-consistency. The second estimator is constructed under conditions ensuring that the least squares criterion can be directly estimated with the parametric rate. These estimators, deeply studied through examples are in particular $\sqrt{n}$-consistent and asymptotically Gaussian in the Cox model and in the excess risk model, whatever is $f\_ε$.

math.ST

Adaptive density deconvolution with dependent inputs

In the convolution model $Z\_i=X\_i+ ε\_i$, we give a model selection procedure to estimate the density of the unobserved variables $(X\_i)\_{1 \leq i \leq n}$, when the sequence $(X\_i)\_{i \geq 1}$ is strictly stationary but not necessarily independent. This procedure depends on wether the density of $ε\_i$ is super smooth or ordinary smooth. The rates of convergence of the penalized contrast estimators are the same as in the independent framework, and are minimax over most classes of regularity on ${\mathbb R}$. Our results apply to mixing sequences, but also to many other dependent sequences. When the errors are super smooth, the condition on the dependence coefficients is the minimal condition of that type ensuring that the sequence $(X\_i)\_{i \geq 1}$ is not a long-memory process.

math.ST

Penalized contrast estimator for adaptive density deconvolution

The authors consider the problem of estimating the density $g$ of independent and identically distributed variables $X\_i$, from a sample $Z\_1, ..., Z\_n$ where $Z\_i=X\_i+σε\_i$, $i=1, ..., n$, $ε$ is a noise independent of $X$, with $σε$ having known distribution. They present a model selection procedure allowing to construct an adaptive estimator of $g$ and to find non-asymptotic bounds for its $\mathbb{L}\_2(\mathbb{R})$-risk. The estimator achieves the minimax rate of convergence, in most cases where lowers bounds are available. A simulation study gives an illustration of the good practical performances of the method.

math.ST

Finite sample penalization in adaptive density deconvolution

We consider the problem of estimating the density $g$ of identically distributed variables $X\_i$, from a sample $Z\_1, ..., Z\_n$ where $Z\_i=X\_i+σε\_i$, $i=1, ..., n$ and $σε\_i$ is a noise independent of $X\_i$ with known density $ σ^{-1}f\_ε(./σ)$. We generalize adaptive estimators, constructed by a model selection procedure, described in Comte et al. (2005). We study numerically their properties in various contexts and we test their robustness. Comparisons are made with respect to deconvolution kernel estimators, misspecification of errors, dependency,... It appears that our estimation algorithm, based on a fast procedure, performs very well in all contexts.

math.ST

Nonparametric Estimation of the Regression Function in an Errors-in-Variables Model

We consider the regression model with errors-in-variables where we observe $n$ i.i.d. copies of $(Y,Z)$ satisfying $Y=f(X)+ξ, Z=X+σε$, involving independent and unobserved random variables $X,ξ,ε$. The density $g$ of $X$ is unknown, whereas the density of $σε$ is completely known. Using the observations $(Y\_i, Z\_i)$, $i=1,...,n$, we propose an estimator of the regression function $f$, built as the ratio of two penalized minimum contrast estimators of $\ell=fg$ and $g$, without any prior knowledge on their smoothness. We prove that its $\mathbb{L}\_2$-risk on a compact set is bounded by the sum of the two $\mathbb{L}\_2(\mathbb{R})$-risks of the estimators of $\ell$ and $g$, and give the rate of convergence of such estimators for various smoothness classes for $\ell$ and $g$, when the errors $ε$ are either ordinary smooth or super smooth. The resulting rate is optimal in a minimax sense in all cases where lower bounds are available.

math.ST