SearcharxivSearch

arXiv subjects

Piet Groeneboom

Publications and source records attributed to Piet Groeneboom.

At least 19 recordsLinked to original sources

Nonparametric Least Squares Estimators for Interval Censoring

The limit distribution of the nonparametric maximum likelihood estimator for interval censored data with more than one observation time per unobservable observation, is still unknown in general. For the so-called separated case, where one has observation times which are at a distance larger than a fixed positive epsilon, the limit distribution was derived in [5]. For the non-separated case there is a conjectured limit distribution, given in [10], Section 5.2 of Part 2. Whether this conjecture holds is still unknown, but the present paper shows that for sample sizes 1000 and 10,000 this limit behavior is still not clearly seen. We prove consistency of a related nonparametric isotonic least squares estimator and sketch of the proof for its limit distribution. We also provide simulation results to show how the nonparametric MLE and least squares estimator behave in comparison. Moreover, we discuss a simpler least squares estimator that can be computed in one step, but is inferior to the other least squares estimator, since it does not use all information. For the simplest model of interval censoring, the current status model, the nonparametric maximum likelihood and least squares estimators are the same. This equivalence breaks down if there are more observation times per unobservable observation. The computations for the simulation of the more complicated interval censoring model were performed by using the iterative convex minorant algorithm. They are provided in the GitHub repository [7].

math.ST

Nonparametric Estimation in Uniform Deconvolution and Interval Censoring

In the uniform deconvolution problem one is interested in estimating the distribution function $F_0$ of a nonnegative random variable, based on a sample with additive uniform noise. A peculiar and not well understood phenomenon of the nonparametric maximum likelihood estimator in this setting is the dichotomy between the situations where $F_0(1)=1$ and $F_0(1)<1$. If $F_0(1)=1$, the MLE can be computed in a straightforward way and its asymptotic pointwise behavior can be derived using the connection to the so-called current status problem. However, if $F_0(1)<1$, one needs an iterative procedure to compute it and the asymptotic pointwise behavior of the nonparametric maximum likelihood estimator is not known. In this paper we describe the problem, connect it to interval censoring problems and a more general model studied in Groeneboom (2024) to state two competing naturally occurring conjectures for the case $F_0(1)<1$. Asymptotic arguments related to smooth functional theory and extensive simulations lead us to to bet on one of these two conjectures.

math.ST

Estimation of the incubation time distribution in the singly and doubly interval censored model

We analyze nonparametric estimators for the distribution function of the incubation time in the singly and doubly interval censoring model. The classical approach is to use parametric families like Weibull, log-normal or gamma distributions in the estimation procedure. We propose nonparametric estimates which stay closer to the data than the classical parametric methods. We also give explicit limit distributions for discrete versions of the models and apply this to compute confidence intervals. The methods complement the analysis of the continuous model. R scripts for computation of the estimates are provided on https://github.com/pietg/incubationtime.

math.ST

Nonparametric estimation of the incubation time distribution

Nonparametric maximum likelihood estimators (MLEs) in inverse problems often have non-normal limit distributions, like Chernoff's distribution. However, if one considers smooth functionals of the model, with corresponding functionals of the MLE, one gets normal limit distributions and faster rates of convergence. We demonstrate this for a model for the incubation time of a disease. The usual approach in the latter models is to use parametric distributions, like Weibull and gamma distributions, which leads to inconsistent estimators. Smoothed bootstrap methods are discussed for constructing confidence intervals. The classical bootstrap, based on the nonparametric MLE itself, has been proved to be inconsistent in this situation.

math.ST

Credible intervals and bootstrap confidence intervals in monotone regression

In the recent paper [5], a Bayesian approach for constructing confidence intervals in monotone regression problems is proposed, based on credible intervals. We view this method from a frequentist point of view, and show that it corresponds to a percentile bootstrap method of which we give two versions. It is shown that a (non-percentile) smoothed bootstrap method has better behavior and does not need correction for over- or undercoverage. The proofs use martingale methods.

math.ST

Confidence intervals in monotone regression

We construct bootstrap confidence intervals for a monotone regression function. It has been shown that the ordinary nonparametric bootstrap, based on the nonparametric least squares estimator (LSE) $\hat f_n$ is inconsistent in this situation. We show, however, that a consistent bootstrap can be based on the smoothed $\hat f_n$, to be called the SLSE (Smoothed Least Squares Estimator). The asymptotic pointwise distribution of the SLSE is derived. The confidence intervals, based on the smoothed bootstrap, are compared to intervals based on the (not necessarily monotone) Nadaraya Watson estimator and the effect of Studentization is investigated. We also give a method for automatic bandwidth choice, correcting work in Sen and Xu (2015). The procedure is illustrated using a well known dataset related to climate change.

math.ST

Nonparametric estimation of the incubation time distribution

We discuss nonparametric estimators of the distribution of the incubation time of a disease. The classical approach in these models is to use parametric families like Weibull, log-normal or gamma in the estimation procedure. We analyze instead the nonparametric maximum likelihood estimator (MLE) and show that, under some conditions, its rate of convergence is cube root $n$ and that its limit behavior is given by Chernoff's distribution. We also study smooth estimates, based on the MLE. The density estimates, based on the MLE, are capable of catching finer or unexpected aspects of the density, in contrast with the classical parametric methods. {\tt R} scripts are provided for the nonparametric methods.

math.ST

Profile least squares estimators in the monotone single index model

We consider least squares estimators of the finite regression parameter $α$ in the single index regression model $Y=ψ(α^T X)+ε$, where $X$ is a $d$-dimensional random vector, $\E(Y|X)=ψ(α^T X)$, and where $ψ$ is monotone. It has been suggested to estimate $α$ by a profile least squares estimator, minimizing $\sum_{i=1}^n(Y_i-ψ(α^T X_i))^2$ over monotone $ψ$ and $α$ on the boundary $S_{d-1}$of the unit ball. Although this suggestion has been around for a long time, it is still unknown whether the estimate is $\sqrt{n}$ convergent. We show that a profile least squares estimator, using the same pointwise least squares estimator for fixed $α$, but using a different global sum of squares, is $\sqrt{n}$-convergent and asymptotically normal. The difference between the corresponding loss functions is studied and also a comparison with other methods is given.

math.ST

Estimation of the incubation time distribution for COVID-19

We consider smooth nonparametric estimation of the incubation time distribution of COVID-19, in connection with the investigation of researchers from the National Institute for Public Health and the Environment (Dutch: RIVM) of 88 travelers from Wuhan: Backer et al (2020). The advantages of the smooth nonparametric approach w.r.t. the parametric approach, using three parametric distributions (Weibull, log-normal and gamma) in Backer et al (2020) is discussed. It is shown that the typical rate of convergence of the smooth estimate of the density is $n^{2/7}$ in a continuous version of the model, where $n$ is the sample size. The (non-smoothed) nonparametric maximum likelihood estimator (MLE) itself is computed by the iterative convex minorant algorithm (Groeneboom and Jongbloed (2014)). All computations are available as {\tt R} scripts in Groeneboom (2020).

stat.AP

Grenander functionals and Cauchy's formula

Let $\hat f_n$ be the nonparametric maximum likelihood estimator of a decreasing density. Grenander characterized this as the left-continuous slope of the least concave majorant of the empirical distribution function. For a sample from the uniform distribution, the asymptotic distribution of the $L_2$-distance of the Grenander estimator to the uniform density was derived in Groeneboom and Pyke (1983) by using a representation of the Grenander estimator in terms of conditioned Poisson and gamma random variables. This representation was also used in Groeneboom and Lopuhaa (1993) to prove a central limit result of Sparre Andersen on the number of jumps of the Grenander estimator. Here we extend this to the proof of a general result on integrals of the Grenander estimator. We also correct Groeneboom and Pyke (1983), where the limit distribution of the sums of gamma and Poisson variables on which the conditioning was done did not have the right form. Saddle point methods and Cauchy's formula are important tools in our development.

math.PR

Elementary Statistics on Trial (the case of Lucia de Berk)

In the conviction of Lucia de Berk an important role was played by a simple hypergeometric model, used by the expert consulted by the court, which produced very small probabilities of occurrences of certain numbers of incidents. We want to draw attention to the fact that, if we take into account the variation among nurses in incidents they experience during their shifts, these probabilities can become considerably larger. This points to the danger of using an oversimplified discrete probability model in these circumstances.

stat.AP

The Lagrange approach in the monotone single index model

The finite-dimensional parameters of the monotone single index model are often estimated by minimization of a least squares criterion and reparametrization to deal with the non-unicity. We avoid the reparametrization by using a Lagrange-type method and replace the minimization over the finite-dimensional parameter alpha by a `crossing of zero' criterion at the derivative level. In particular, we consider a simple score estimator (SSE), an efficient score estimator (ESE), and a penalized least squares estimator (PLSE) for which we can apply this method. The SSE and ESE were discussed in Balabdaoui, Groeneboom and Hendrickx (2018}, but the proofs still used reparametrization. Another version of the PLSE was discussed in Kuchibhotla and Patra (2017), where also reparametrization was used. The estimators are compared with the profile least squares estimator (LSE), Han's maximum rank estimator (MRE), the effective dimension reduction estimator (EDR) and a linear least squares estimator, which can be used if the covariates have an elliptically symmetric distribution. We also investigate the effects of random starting values in the search algorithms.

stat.CO

Score estimation in the monotone single index model

We consider estimation in the single index model where the link function is monotone. For this model a profile least squares estimator has been proposed to estimate the unknown link function and index. Although it is natural to propose this procedure, it is still unknown whether it produces index estimates which converge at the parametric rate. We show that this holds if we solve a score equation corresponding to this least squares problem. Using a Lagrangian formulation, we show how one can solve this score equation without any reparametrization. This makes it easy to solve the score equations in high dimensions. We also compare our method with the Effective Dimension Reduction (EDR) and the Penalized Least Squares Estimator (PLSE) methods, both available on CRAN as R packages, and compare with link-free methods, where the covariates are ellipticallly symmetric.

math.ST

The nonparametric bootstrap for the current status model

It has been proved that direct bootstrapping of the nonparametric maximum likelihood estimator (MLE) of the distribution function in the current status model leads to inconsistent confidence intervals. We show that bootstrapping of functionals of the MLE can however be used to produce valid intervals. To this end, we prove that the bootstrapped MLE converges at the right rate in the $L_p$-distance. We also discuss applications of this result to the current status regression model.

stat.ME

Current status linear regression

We construct $\sqrt{n}$-consistent and asymptotically normal estimates for the finite dimensional regression parameter in the current status linear regression model, which do not require any smoothing device and are based on maximum likelihood estimates (MLEs) of the infinite dimensional parameter. We also construct estimates, again only based on these MLEs, which are arbitrarily close to efficient estimates, if the generalized Fisher information is finite. This type of efficiency is also derived under minimal conditions for estimates based on smooth non-monotone plug-in estimates of the distribution function. Algorithms for computing the estimates and for selecting the bandwidth of the smooth estimates with a bootstrap method are provided. The connection with results in the econometric literature is also pointed out.

math.ST

Confidence intervals for the current status model

We discuss a new way of constructing pointwise confidence intervals for the distribution function in the current status model. The confidence intervals are based on the smoothed maximum likelihood estimator (SMLE) and constructed using bootstrap methods. Other methods to construct confidence intervals, using the non-standard limit distribution of the (restricted) MLE, are compared to our approach via simulations and real data applications.

math.ST

Nonparametric confidence intervals for monotone functions

We study nonparametric isotonic confidence intervals for monotone functions. In Banerjee and Wellner (2001) pointwise confidence intervals, based on likelihood ratio tests for the restricted and unrestricted MLE in the current status model, are introduced. We extend the method to the treatment of other models with monotone functions, and demonstrate our method by a new proof of the results in Banerjee and Wellner (2001) and also by constructing confidence intervals for monotone densities, for which still theory had to be developed. For the latter model we prove that the limit distribution of the LR test under the null hypothesis is the same as in the current status model. We compare the confidence intervals, so obtained, with confidence intervals using the smoothed maximum likelihood estimator (SMLE), using bootstrap methods. The `Lagrange-modified' cusum diagrams, developed here, are an essential tool both for the computation of the restricted MLEs and for the development of the theory for the confidence intervals, based on the LR tests.

math.ST

Maximum smoothed likelihood estimators for the interval censoring model

We study the maximum smoothed likelihood estimator (MSLE) for interval censoring, case 2, in the so-called separated case. Characterizations in terms of convex duality conditions are given and strong consistency is proved. Moreover, we show that, under smoothness conditions on the underlying distributions and using the usual bandwidth choice in density estimation, the local convergence rate is $n^{-2/5}$ and the limit distribution is normal, in contrast with the rate $n^{-1/3}$ of the ordinary maximum likelihood estimator.

math.ST