SearcharxivSearch

arXiv subjects

Estelle Kuhn

Publications and source records attributed to Estelle Kuhn.

15 recordsLinked to original sources

Estimation and variable selection in high dimension in a causal joint model of survival times and longitudinal outcomes with random effects

We consider a joint survival and mixed-effects model to explain the survival time from longitudinal data and high-dimensional covariates in a population. The longitudinal data is modeled using a non linear mixed-effects model to account for the inter-individual variability in the population. The corresponding regression function serves as a link function incorporated into the survival model. In that way, the longitudinal data is related to the survival time. We consider a Cox model that takes into account both high-dimensional covariates and the link function. There are two main objectives: first, identify the relevant covariates that contribute to explaining survival time, and second, estimate all unknown parameters of the joint model. For the first objective, we consider the estimate defined by maximizing the marginal log-likelihood regularized with a l1-penalty term. To tackle the optimization problem, we implement an adaptive stochastic gradient to handle the latent variables of the non linear mixed-effects model associated with a proximal operator to manage the non-differentiability of the penalty. We rely on an eBIC model choice criterion to select an optimal value for the regularization parameter. Once the relevant covariates are selected, we re-estimate the parameters in the reduced model by maximizing the likelihood using an adaptive stochastic gradient descent. We provide relevant simulations that showcase the performance of the proposed variable selection and parameter estimation method in the joint model. We investigate the effect of censoring and of the presence of correlation between the individual parameters in the mixed model.

math.ST

Estimation and variable selection in high dimension in nonlinear mixed-effects models

We consider nonlinear mixed effects models including high-dimensional covariates to model individual parameters variability. The objective is to identify relevant covariates among a large set under sparsity assumption and to estimate model parameters. To face the high dimensional setting we consider a regularized estimator namely the maximum likelihood estimator penalized with the l1-penalty. We rely on the use of the eBIC model choice criteria to select an optimal reduced model. Then we estimate the parameters by maximizing the likelihood of the reduced model. We calculate in practice the maximum likelihood estimator penalized with the l1-penalty though a weighted proximal stochastic gradient descent algorithm with an adaptive learning rate. This choice allows us to consider very general models, in particular models that do not belong to the curved exponential family. We demonstrate first in a simple linear toy model through a simulation study the good convergence properties of this optimization algorithm. We compare then the performance of the proposed methodology with those of the \glmmLasso procedure in a linear mixed effects model in a simulation study. We illustrate also its performance in a nonlinear mixed-effects logistic growth model through simulation. We finally highlight the beneficit of the proposed procedure relying on an integrated single step approach regarding two others two steps approaches for variable selection objective.

math.ST

Estimation of ratios of normalizing constants using stochastic approximation : the SARIS algorithm

Computing ratios of normalizing constants plays an important role in statistical modeling. Two important examples are hypothesis testing in latent variables models, and model comparison in Bayesian statistics. In both examples, the likelihood ratio and the Bayes factor are defined as the ratio of the normalizing constants of posterior distributions. We propose in this article a novel methodology that estimates this ratio using stochastic approximation principle. Our estimator is consistent and asymptotically Gaussian. Its asymptotic variance is smaller than the one of the popular optimal bridge sampling estimator. Furthermore, it is much more robust to little overlap between the two unnormalized distributions considered. Thanks to its online definition, our procedure can be integrated in an estimation process in latent variables model, and therefore reduce the computational effort. The performances of the estimator are illustrated through a simulation study and compared to two other estimators : the ratio importance sampling and the optimal bridge sampling estimators.

stat.AP

Bootstrap test procedure for variance components in nonlinear mixed effects models in the presence of nuisance parameters and a singular Fisher Information Matrix

We examine the problem of variance components testing in general mixed effects models using the likelihood ratio test. We account for the presence of nuisance parameters, i.e. the fact that some untested variances might also be equal to zero. Two main issues arise in this context leading to a non regular setting. First, under the null hypothesis the true parameter value lies on the boundary of the parameter space. Moreover, due to the presence of nuisance parameters the exact location of these boundary points is not known, which prevents from using classical asymptotic theory of maximum likelihood estimation. Then, in the specific context of nonlinear mixed-effects models, the Fisher information matrix is singular at the true parameter value. We address these two points by proposing a shrinked parametric bootstrap procedure, which is straightforward to apply even for nonlinear models. We show that the procedure is consistent, solving both the boundary and the singularity issues, and we provide a verifiable criterion for the applicability of our theoretical results. We show through a simulation study that, compared to the asymptotic approach, our procedure has a better small sample performance and is more robust to the presence of nuisance parameters. A real data application is also provided.

stat.ME

Efficient preconditioned stochastic gradient descent for estimation in latent variable models

Latent variable models are powerful tools for modeling complex phenomena involving in particular partially observed data, unobserved variables or underlying complex unknown structures. Inference is often difficult due to the latent structure of the model. To deal with parameter estimation in the presence of latent variables, well-known efficient methods exist, such as gradient-based and EM-type algorithms, but with practical and theoretical limitations. In this paper, we propose as an alternative for parameter estimation an efficient preconditioned stochastic gradient algorithm. Our method includes a preconditioning step based on a positive definite Fisher information matrix estimate. We prove convergence results for the proposed algorithm under mild assumptions for very general latent variables models. We illustrate through relevant simulations the performance of the proposed methodology in a nonlinear mixed effects model and in a stochastic block model.

math.ST

Estimating Fisher Information Matrix in Latent Variable Models based on the Score Function

The Fisher information matrix (FIM) is a key quantity in statistics as it is required for example for evaluating asymptotic precisions of parameter estimates, for computing test statistics or asymptotic distributions in statistical testing, for evaluating post model selection inference results or optimality criteria in experimental designs. However its exact computation is often not trivial. In particular in many latent variable models, it is intricated due to the presence of unobserved variables. Therefore the observed FIM is usually considered in this context to estimate the FIM. Several methods have been proposed to approximate the observed FIM when it can not be evaluated analytically. Among the most frequently used approaches are Monte-Carlo methods or iterative algorithms derived from the missing information principle. All these methods require to compute second derivatives of the complete data log-likelihood which leads to some disadvantages from a computational point of view. In this paper, we present a new approach to estimate the FIM in latent variable model. The advantage of our method is that only the first derivatives of the log-likelihood is needed, contrary to other approaches based on the observed FIM. Indeed we consider the empirical estimate of the covariance matrix of the score. We prove that this estimate of the Fisher information matrix is unbiased, consistent and asymptotically Gaussian. Moreover we highlight that none of both estimates is better than the other in terms of asymptotic covariance matrix. When the proposed estimate can not be directly analytically evaluated, we present a stochastic approximation estimation algorithm to compute it. This algorithm provides this estimate of the FIM as a by-product of the parameter estimates. We emphasize that the proposed algorithm only requires to compute the first derivatives of the complete data log-likelihood with respect to the parameters. We prove that the estimation algorithm is consistent and asymptotically Gaussian when the number of iterations goes to infinity. We evaluate the finite sample size properties of the proposed estimate and of the observed FIM through simulation studies in linear mixed effects models and mixture models. We also investigate the convergence properties of the estimation algorithm in non linear mixed effects models. We compare the performances of the proposed algorithm to those of other existing methods.

stat.ME

varTestnlme: an R package for Variance Components Testing in Linear and Nonlinear Mixed-effects Models

The issue of variance components testing arises naturally when building mixed-effects models, to decide which effects should be modeled as fixed or random. While tests for fixed effects are available in R for models fitted with lme4, tools are missing when it comes to random effects. The varTestnlme package for R aims at filling this gap. It allows to test whether any subset of the variances and covariances corresponding to any subset of the random effects, are equal to zero using asymptotic property of the likelihood ratio test statistic. It also offers the possibility to test simultaneously for fixed effects and variance components. It can be used for linear, generalized linear or nonlinear mixed-effects models fitted via lme4, nlme or saemix. Theoretical properties of the used likelihood ratio test are recalled, numerical methods used to implement the test procedure are detailed and examples based on different real datasets using different mixed models are provided.

stat.ME

Modeling dependent survival data through random effects with spatial correlation at the subject level

Dynamical phenomena such as infectious diseases are often investigated by following up subjects longitudinally, thus generating time to event data. The spatial aspect of such data is also of primordial importance, as many infectious diseases are transmitted from one subject to another. In this paper, a spatially correlated frailty model is introduced that accommodates for the correlation between subjects based on the distance between them. Estimates are obtained through a stochastic approximation version of the Expectation Maximization algorithm combined with a Monte-Carlo Markov Chain, for which convergence is proven. The novelty of this model is that spatial correlation is introduced for survival data at the subject level, each subject having its own frailty. This univariate spatially correlated frailty model is used to analyze spatially dependent malaria data, and its results are compared with other standard models.

stat.ME

Properties of the Stochastic Approximation EM Algorithm with Mini-batch Sampling

To deal with very large datasets a mini-batch version of the Monte Carlo Markov Chain Stochastic Approximation Expectation-Maximization algorithm for general latent variable models is proposed. For exponential models the algorithm is shown to be convergent under classicalconditions as the number of iterations increases. Numerical experiments illustrate the performance of the mini-batch algorithm in various models.In particular, we highlight that mini-batch sampling results in an important speed-up of the convergence of the sequence of estimators generated by the algorithm. Moreover, insights on the effect of the mini-batch size on the limit distribution are presented. Finally, we illustrate how to use mini-batch sampling in practice to improve results when a constraint on the computing time is given.

stat.CO

Convergent stochastic algorithm for parameter estimation in frailty models using integrated partial likelihood

Frailty models are often the model of choice for heterogeneous survival data. A frailty model contains both random effects and fixed effects, with the random effects accommodating for the correlation in the data. Different estimation procedures have been proposed for the fixed effects and the variances of and covariances between the random effects. Especially with an unspecified baseline hazard, i.e., the Cox model, the few available methods deal only with a specific correlation structure. In this paper, an estimation procedure, based on the integrated partial likelihood, is introduced, which can generally deal with any kind of correlation structure. The new approach, namely the maximisation of the integrated partial likelihood, combined with a stochastic estimation procedure allows also for a wide choice of distributions for the random effects. First, we demonstrate the almost sure convergence of the stochastic algorithm towards a critical point of the integrated partial likelihood. Second, numerical convergence properties are evaluated by simulation. Third, the advantage of using an unspecified baseline hazard is demonstrated through application on cancer clinical trial data.

stat.ME

Likelihood ratio test for variance components in nonlinear mixed effects models

Mixed effects models are widely used to describe heterogeneity in a population. A crucial issue when adjusting such a model to data consists in identifying fixed and random effects. From a statistical point of view, it remains to test the nullity of the variances of a given subset of random effects. Some authors have proposed to use the likelihood ratio test and have established its asymptotic distribution in some particular cases. Nevertheless, to the best of our knowledge, no general variance components testing procedure has been fully investigated yet. In this paper, we study the likelihood ratio test properties to test that the variances of a general subset of the random effects are equal to zero in both linear and nonlinear mixed effects model, extending the existing results. We prove that the asymptotic distribution of the test is a chi-bar-square distribution, that is to say a mixture of chi-square distributions, and we identify the corresponding weights. We highlight in particular that the limiting distribution depends on the presence of correlations between the random effects but not on the linear or nonlinear structure of the mixed effects model. We illustrate the finite sample size properties of the test procedure through simulation studies and apply the test procedure to two real datasets of dental growth and of coucal growth.

stat.ME

Convergence of the Wang-Landau algorithm

We analyze the convergence properties of the Wang-Landau algorithm. This sampling method belongs to the general class of adaptive importance sampling strategies which use the free energy along a chosen reaction coordinate as a bias. Such algorithms are very helpful to enhance the sampling properties of Markov Chain Monte Carlo algorithms, when the dynamics is metastable. We prove the convergence of the Wang-Landau algorithm and an associated central limit theorem.

math.PR

Convergent Stochastic Expectation Maximization algorithm with efficient sampling in high dimension. Application to deformable template model estimation

Estimation in the deformable template model is a big challenge in image analysis. The issue is to estimate an atlas of a population. This atlas contains a template and the corresponding geometrical variability of the observed shapes. The goal is to propose an accurate algorithm with low computational cost and with theoretical guaranties of relevance. This becomes very demanding when dealing with high dimensional data which is particularly the case of medical images. We propose to use an optimized Monte Carlo Markov Chain method into a stochastic Expectation Maximization algorithm in order to estimate the model parameters by maximizing the likelihood. In this paper, we present a new Anisotropic Metropolis Adjusted Langevin Algorithm which we use as transition in the MCMC method. We first prove that this new sampler leads to a geometrically uniformly ergodic Markov chain. We prove also that under mild conditions, the estimated parameters converge almost surely and are asymptotically Gaussian distributed. The methodology developed is then tested on handwritten digits and some 2D and 3D medical images for the deformable model estimation. More widely, the proposed algorithm can be used for a large range of models in many fields of applications such as pharmacology or genetic.

math.ST

Construction of Bayesian Deformable Models via Stochastic Approximation Algorithm: A Convergence Study

The problem of the definition and the estimation of generative models based on deformable templates from raw data is of particular importance for modelling non aligned data affected by various types of geometrical variability. This is especially true in shape modelling in the computer vision community or in probabilistic atlas building for Computational Anatomy (CA). A first coherent statistical framework modelling the geometrical variability as hidden variables has been given by Allassonnière, Amit and Trouvé (JRSS 2006). Setting the problem in a Bayesian context they proved the consistency of the MAP estimator and provided a simple iterative deterministic algorithm with an EM flavour leading to some reasonable approximations of the MAP estimator under low noise conditions. In this paper we present a stochastic algorithm for approximating the MAP estimator in the spirit of the SAEM algorithm. We prove its convergence to a critical point of the observed likelihood with an illustration on images of handwritten digits.

stat.CO

Stochastic Algorithm For Parameter Estimation For Dense Deformable Template Mixture Model

Estimating probabilistic deformable template models is a new approach in the fields of computer vision and probabilistic atlases in computational anatomy. A first coherent statistical framework modelling the variability as a hidden random variable has been given by Allassonnière, Amit and Trouvé in [1] in simple and mixture of deformable template models. A consistent stochastic algorithm has been introduced in [2] to face the problem encountered in [1] for the convergence of the estimation algorithm for the one component model in the presence of noise. We propose here to go on in this direction of using some "SAEM-like" algorithm to approximate the MAP estimator in the general Bayesian setting of mixture of deformable template model. We also prove the convergence of this algorithm toward a critical point of the penalised likelihood of the observations and illustrate this with handwritten digit images.

stat.CO