SearcharxivSearch

arXiv subjects

Marco Alfo'

Publications and source records attributed to Marco Alfo'.

5 recordsLinked to original sources

Omitted covariates bias and finite mixtures of regression models for longitudinal responses

Individual-specific, time-constant, random effects are often used to model dependence and/or to account for omitted covariates in regression models for longitudinal responses. Longitudinal studies have known a huge and widespread use in the last few years as they allow to distinguish between so-called age and cohort effects; these relate to differences that can be observed at the beginning of the study and stay persistent through time, and changes in the response that are due to the temporal dynamics in the observed covariates. While there is a clear and general agreement on this purpose, the random effect approach has been frequently criticized for not being robust to the presence of correlation between the observed (i.e. covariates) and the unobserved (i.e. random effects) heterogeneity. Starting from the so-called correlated effect approach, we argue that the random effect approach may be parametrized to account for potential correlation between observables and unobservables. Specifically, when the random effect distribution is estimated non-parametrically using a discrete distribution on finite number of locations, a further, more general, solution is developed. This is illustrated via a large scale simulation study and the analysis of a benchmark dataset.

stat.ME

Semi-Parametric Empirical Best Prediction for small area estimation of unemployment indicators

The Italian National Institute for Statistics regularly provides estimates of unemployment indicators using data from the Labor Force Survey. However, direct estimates of unemployment incidence cannot be released for Local Labor Market Areas. These are unplanned domains defined as clusters of municipalities; many are out-of-sample areas and the majority is characterized by a small sample size, which render direct estimates inadequate. The Empirical Best Predictor represents an appropriate, model-based, alternative. However, for non-Gaussian responses, its computation and the computation of the analytic approximation to its Mean Squared Error require the solution of (possibly) multiple integrals that, generally, have not a closed form. To solve the issue, Monte Carlo methods and parametric bootstrap are common choices, even though the computational burden is a non trivial task. In this paper, we propose a Semi-Parametric Empirical Best Predictor for a (possibly) non-linear mixed effect model by leaving the distribution of the area-specific random effects unspecified and estimating it from the observed data. This approach is known to lead to a discrete mixing distribution which helps avoid unverifiable parametric assumptions and heavy integral approximations. We also derive a second-order, bias-corrected, analytic approximation to the corresponding Mean Squared Error. Finite sample properties of the proposed approach are tested via a large scale simulation study. Furthermore, the proposal is applied to unit-level data from the 2012 Italian Labor Force Survey to estimate unemployment incidence for 611 Local Labor Market Areas using auxiliary information from administrative registers and the 2011 Census.

stat.ME

A non-homogeneous hidden Markov model for partially observed longitudinal responses

Dropout represents a typical issue to be addressed when dealing with longitudinal studies. If the mechanism leading to missing information is non-ignorable, inference based on the observed data only may be severely biased. A frequent strategy to obtain reliable parameter estimates is based on the use of individual-specific random coefficients that help capture sources of unobserved heterogeneity and, at the same time, define a reasonable structure of dependence between the longitudinal and the missing data process. We refer to elements in this class as random coefficient based dropout models (RCBDMs). We propose a dynamic, semi-parametric, version of the standard RCBDM to deal with discrete time to event. Time-varying random coefficients that evolve over time according to a non-homogeneous hidden Markov chain are considered to model dependence between longitudinal responses recorded from the same subject. A separate set of random coefficients is considered to model dependence between missing data indicators. Last, the joint distribution of the random coefficients in the two equations helps describe the dependence between the two processes. To ensure model flexibility and avoid unverifiable assumptions, we leave the joint distribution of the random coefficients unspecified and estimate it via nonparametric maximum likelihood. The proposal is applied to data from the Leiden 85+ study on the evolution of cognitive functioning in the elderly.

stat.ME

M-quantile regression for multivariate longitudinal data: analysis of the Millennium Cohort Study data

We propose a M-quantile regression model for the analysis of multivariate, continuous, longitudinal data. M-quantile regression represents an appealing alternative to standard regression models, as it combines the robustness of quantile and the efficiency of expectile regression, providing a complete picture of the response variable distribution. Discrete, individual-specific, random parameters are used to account for both dependence within the same response recorded at different times and association between different responses observed on the same sample unit at a given time. A suitable parametrisation is also introduced in the linear predictor to account for possible dependence between the individual specific random parameters and the vector of observed covariates, that is to account for endogeneity of some covariates. An extended EM algorithm is proposed to derive model parameter estimates under a maximum likelihood approach. The model is applied to the analysis of the strengths and difficulties questionnaire scores from the Millennium Cohort Study in the UK.

stat.ME

Quantile regression for longitudinal data: unobserved heterogeneity and informative missingness

Linear quantile regression models aim at providing a detailed and robust picture of the (conditional) response distribution as function of a set of observed covariates. Longitudinal data represent an interesting field of application of such models; due to their peculiar features, they represent a substantial challenge, in that the standard, cross-sectional, model representation needs to be extended for dealing with such kind of data. In fact, repeated observations from the same statistical unit poses a problem of dependence; in a conditional perspective, this dependence could be ascribed to sources of unobserved, individual-specific, heterogeneity. Along these lines, quantile regression models have recently been extended to the analysis of longitudinal, continuous, responses, by modelling dependence via time-constant or time-varying random effects. In this manuscript, we introduce a general quantile regression model for longitudinal, continuous, responses where time-varying and time-constant random parameters are jointly taken into account. A further feature of longitudinal designs is the presence of partially incomplete sequences, due to some individuals leaving the study before its designed end. The missing data process may produce a selection of units which can be informative with respect to the parameters of the longitudinal data model. To deal with the case of irretrievable drop-out, we introduce a pattern mixture version of the linear quantile hidden Markov model, where we account for time-varying heterogeneity and for changes in the fixed effect vector due to differential propensities to stay in the study. The proposed models are illustrated using a well known benchmark dataset on longitudinal dynamics of CD4 cells and by means of a large scale simulation study, entailing different quantiles and both complete and partially complete (ie subject to drop-out) individual sequences.

stat.ME