SearcharxivSearch

arXiv subjects

Karine Bertin

Publications and source records attributed to Karine Bertin.

At least 19 recordsLinked to original sources

Adaptive estimation in regression models for weakly dependent data and explanatory variable with known density

This article is dedicated to the estimation of the regression function when the explanatory variable is a weakly dependent process whose correlation coefficient exhibits exponential decay and has a known bounded density function. The accuracy of the estimation is measured using pointwise risk. A data-driven procedure is proposed using kernel estimation with bandwidth selected via the Goldenshluger-Lepski approach. We demonstrate that the resulting estimator satisfies an oracle-type inequality and it is also shown to be adaptive over Hölder classes. Additionally, unsupervised statistical learning techniques are described and applied to calibrate the method, and some simulations are provided to illustrate the performance of the method.

math.ST

A new adaptive local polynomial density estimation procedure on complicated domains

This paper presents a novel approach for pointwise estimation of multivariate density functions on known domains of arbitrary dimensions using nonparametric local polynomial estimators. Our method is highly flexible, as it applies to both simple domains, such as open connected sets, and more complicated domains that are not star-shaped around the point of estimation. This enables us to handle domains with sharp concavities, holes, and local pinches, such as polynomial sectors. Additionally, we introduce a data-driven selection rule based on the general ideas of Goldenshluger and Lepski. Our results demonstrate that the local polynomial estimators are minimax under a $L^2$ risk across a wide range of Hölder-type functional classes. In the adaptive case, we provide oracle inequalities and explicitly determine the convergence rate of our statistical procedure. Simulations on polynomial sectors show that our oracle estimates outperform those of the most popular alternative method, found in the sparr package for the R software. Our statistical procedure is implemented in an online R package which is readily accessible.

math.ST

Minimax properties of Dirichlet kernel density estimators

This paper considers the asymptotic behavior in $β$-Hölder spaces, and under $L^p$ losses, of a Dirichlet kernel density estimator proposed by Aitchison and Lauder (1985) for the analysis of compositional data. In recent work, Ouimet and Tolosana-Delgado (2022) established the uniform strong consistency and asymptotic normality of this estimator. As a complement, it is shown here that the Aitchison-Lauder estimator can achieve the minimax rate asymptotically for a suitable choice of bandwidth whenever $(p,β) \in [1, 3) \times (0, 2]$ or $(p, β) \in \mathcal{A}_d$, where $\mathcal{A}_d$ is a specific subset of $[3, 4) \times (0, 2]$ that depends on the dimension $d$ of the Dirichlet kernel. It is also shown that this estimator cannot be minimax when either $p \in [4, \infty)$ or $β\in (2, \infty)$. These results extend to the multivariate case, and also rectify in a minor way, earlier findings of Bertin and Klutchnikoff (2011) concerning the minimax properties of Beta kernel estimators.

math.ST

Least square estimators in linear regression models under negatively superadditive dependent random observations

In this article we study the asymptotic behaviour of the least square estimator in a linear regression model based on random observation instances. We provide mild assumptions on the moments and dependence structure on the randomly spaced observations and the residuals under which the estimator is strongly consistent. In particular, we consider observation instances that are negatively superadditive dependent within each other, while for the residuals we merely assume that they are generated by some continuous function. In addition, we prove that the rate of convergence is proportional to the sampling rate $N$, and we complement our findings with a simulation study providing insights on finite sample properties.

math.ST

Adaptive regression with Brownian path covariate

This paper deals with estimation with functional covariates. More precisely, we aim at estimating the regression function $m$ of a continuous outcome $Y$ against a standard Wiener coprocess $W$. Following Cadre and Truquet (2015) and Cadre, Klutchnikoff, and Massiot (2017) the Wiener-Itô decomposition of $m(W)$ is used to construct a family of estimators. The minimax rate of convergence over specific smoothness classes is obtained. A data-driven selection procedure is defined following the ideas developed by Goldenshluger and Lepski (2011). An oracle-type inequality is obtained which leads to adaptive results.

math.ST

High dimensional VAR with low rank transition

We propose a vector auto-regressive (VAR) model with a low-rank constraint on the transition matrix. This new model is well suited to predict high-dimensional series that are highly correlated, or that are driven by a small number of hidden factors. We study estimation, prediction, and rank selection for this model in a very general setting. Our method shows excellent performances on a wide variety of simulated datasets. On macro-economic data from Giannone et al. (2015), our method is competitive with state-of-the-art methods in small dimension, and even improves on them in high dimension.

math.ST

AutoRegressive Planet Search: Application to the Kepler Mission

The 4-year light curves of 156,717 stars observed with NASA's Kepler mission are analyzed using the AutoRegressive Planet Search (ARPS) methodology described by Caceres et al. (2019). The three stages of processing are: maximum likelihood ARIMA modeling of the light curves to reduce stellar brightness variations; constructing the Transit Comb Filter periodogram to identify transit-like periodic dips in the ARIMA residuals; Random Forest classification trained on Kepler Team confirmed planets using several dozen features from the analysis. Orbital periods between 0.2 and 100 days are examined. The result is a recovery of 76% of confirmed planets, 97% when period and transit depth constraints are added. The classifier is then applied to the full Kepler dataset; 1,004 previously noticed and 97 new stars have light curve criteria consistent with the confirmed planets, after subjective vetting removes clear False Alarms and False Positive cases. The 97 Kepler ARPS Candidate Transits mostly have periods $P<10$ days; many are UltraShort Period hot planets with radii $<1$% of the host star. Extensive tabular and graphical output from the ARPS time series analysis is provided to assist in other research relating to the Kepler sample.

astro-ph.EP

AutoRegressive Planet Search: Methodology

The detection of periodic signals from transiting exoplanets is often impeded by extraneous aperiodic photometric variability, either intrinsic to the star or arising from the measurement process. Frequently, these variations are autocorrelated wherein later flux values are correlated with previous ones. In this work, we present the methodology of the Autoregessive Planet Search (ARPS) project which uses Autoregressive Integrated Moving Average (ARIMA) and related statistical models that treat a wide variety of stochastic processes, as well as nonstationarity, to improve detection of new planetary transits. Providing a time series is evenly spaced or can be placed on an evenly spaced grid with missing values, these low-dimensional parametric models can prove very effective. We introduce a planet-search algorithm to detect periodic transits in the residuals after the application of ARIMA models. Our matched-filter algorithm, the Transit Comb Filter (TCF), is closely related to the traditional Box-fitting Least Squares and provides an analogous periodogram. Finally, if a previously identified or simulated sample of planets is available, selected scalar features from different stages of the analysis -- the original light curves, ARIMA fits, TCF periodograms, and folded light curves -- can be collectively used with a multivariate classifier to identify promising candidates while efficiently rejecting false alarms. We use Random Forests for this task, in conjunction with Receiver Operating Characteristic (ROC) curves, to define discovery criteria for new, high fidelity planetary candidates. The ARPS methodology can be applied to both evenly spaced satellite light curves and densely cadenced ground-based photometric surveys.

astro-ph.EP

Adaptive Density Estimation on Bounded Domains

We study the estimation, in Lp-norm, of density functions defined on [0,1]^d. We construct a new family of kernel density estimators that do not suffer from the so-called boundary bias problem and we propose a data-driven procedure based on the Goldenshluger and Lepski approach that jointly selects a kernel and a bandwidth. We derive two estimators that satisfy oracle-type inequalities. They are also proved to be adaptive over a scale of anisotropic or isotropic Sobolev-Slobodetskii classes (which are particular cases of Besov or Sobolev classical classes). The main interest of the isotropic procedure is to obtain adaptive results without any restriction on the smoothness parameter.

math.ST

A Bayesian approach for the segmentation of series corrupted by a functional part

We propose a Bayesian approach to detect multiple change-points in a piecewise-constant signal corrupted by a functional part corresponding to environmental or experimental disturbances. The piecewise constant part (also called segmentation part) is expressed as the product of a lower triangular matrix by a sparse vector. The functional part is a linear combination of functions from a large dictionary. A Stochastic Search Variable Selection approach is used to obtain sparse estimations of the segmentation parameters (the change-points and the means over the segments) and of the functional part. The performance of our proposed method is assessed using simulation experiments. Applications to two real datasets from geodesy and economy fields are also presented.

math.ST

Pointwise Adaptive Estimation of the MarginalDensity of a Weakly Dependent Process

This paper is devoted to the estimation of the common marginal density function of weakly dependent processes. The accuracy of estimation is measured using pointwise risks. We propose a datadriven procedure using kernel rules. The bandwidth is selected using the approach of Goldenshluger and Lepski and we prove that the resulting estimator satisfies an oracle type inequality. The procedure is also proved to be adaptive (in a minimax framework) over a scale of Hölder balls for several types of dependence: stong mixing processes, $λ$-dependent processes or i.i.d. sequences can be considered using a single procedure of estimation. Some simulations illustrate the performance of the proposed method.

math.ST

Adaptive pointwise estimation of conditional density function

In this paper we consider the problem of estimating $f$, the conditional density of $Y$ given $X$, by using an independent sample distributed as $(X,Y)$ in the multivariate setting. We consider the estimation of $f(x,.)$ where $x$ is a fixed point. We define two different procedures of estimation, the first one using kernel rules, the second one inspired from projection methods. Both adapted estimators are tuned by using the Goldenshluger and Lepski methodology. After deriving lower bounds, we show that these procedures satisfy oracle inequalities and are optimal from the minimax point of view on anisotropic H{ö}lder balls. Furthermore, our results allow us to measure precisely the influence of $\mathrm{f}\_X(x)$ on rates of convergence, where $\mathrm{f}\_X$ is the density of $X$. Finally, some simulations illustrate the good behavior of our tuned estimates in practice.

math.ST

Segmentation of multiple series using a Lasso strategy

We propose a new semi-parametric approach to the joint segmentation of multiple series corrupted by a functional part. This problem appears in particular in geodesy where GPS permanent station coordinate series are affected by undocumented artificial abrupt changes and additionally show prominent periodic variations. Detecting and estimating them are crucial, since those series are used to determine averaged reference coordinates in geosciences and to infer small tectonic motions induced by climate change. We propose an iterative procedure based on Dynamic Programming for the segmentation part and Lasso estimators for the functional part. Our Lasso procedure, based on the dictionary approach, allows us to both estimate smooth functions and functions with local irregularity, which permits more flexibility than previous proposed methods. This yields to a better estimation of the bias part and improvements in the segmentation. The performance of our method is assessed using simulated and real data. In particular, we apply our method to data from four GPS stations in Yarragadee, Australia. Our estimation procedure results to be a reliable tool to assess series in terms of change detection and periodic variations estimation giving an interpretable estimation of the functional part of the model in terms of known functions.

stat.ME

Lasso-type estimators for Semiparametric Nonlinear Mixed-Effects Models Estimation

Parametric nonlinear mixed effects models (NLMEs) are now widely used in biometrical studies, especially in pharmacokinetics research and HIV dynamics models, due to, among other aspects, the computational advances achieved during the last years. However, this kind of models may not be flexible enough for complex longitudinal data analysis. Semiparametric NLMEs (SNMMs) have been proposed by Ke and Wang (2001). These models are a good compromise and retain nice features of both parametric and nonparametric models resulting in more flexible models than standard parametric NLMEs. However, SNMMs are complex models for which estimation still remains a challenge. The estimation procedure proposed by Ke and Wang (2001) is based on a combination of log-likelihood approximation methods for parametric estimation and smoothing splines techniques for nonparametric estimation. In this work, we propose new estimation strategies in SNMMs. On the one hand, we use the Stochastic Approximation version of EM algorithm (Delyon et al., 1999) to obtain exact ML and REML estimates of the fixed effects and variance components. On the other hand, we propose a LASSO-type method to estimate the unknown nonlinear function. We derive oracle inequalities for this nonparametric estimator. We combine the two approaches in a general estimation procedure that we illustrate with simulated and real data.

stat.ME

Minimax properties of beta kernel density estimators

In this paper, we are interested in the study of beta kernel estimators from an asymptotic minimax point of view. It is well known that beta kernel estimators are, on the contrary of classical kernel estimators, "free of boundary effect" and thus are very useful in practice. The goal of this paper is to prove that there is a price to pay: for very regular functions or for certain losses, these estimators are not minimax. Nevertheless they are minimax for classical regularities such as regularity of order two or less than two, supposed commonly in the practice and for some classical losses.

math.ST

Maximum likelihood estimators and random walks in long memory models

We consider statistical models driven by Gaussian and non-Gaussian self-similar processes with long memory and we construct maximum likelihood estimators (MLE) for the drift parameter. Our approach is based on the approximation by random walks of the driving noise. We study the asymptotic behavior of the estimators and we give some numerical simulations to illustrate our results.

math.ST

Adaptive Dantzig density estimation

This paper deals with the problem of density estimation. We aim at building an estimate of an unknown density as a linear combination of functions of a dictionary. Inspired by Candès and Tao's approach, we propose an $\ell_1$-minimization under an adaptive Dantzig constraint coming from sharp concentration inequalities. This allows to consider a wide class of dictionaries. Under local or global coherence assumptions, oracle inequalities are derived. These theoretical results are also proved to be valid for the natural Lasso estimate associated with our Dantzig procedure. Then, the issue of calibrating these procedures is studied from both theoretical and practical points of view. Finally, a numerical study shows the significant improvement obtained by our procedures when compared with other classical procedures.

math.ST