SearcharxivSearch

arXiv subjects

Nicola Salvati

Publications and source records attributed to Nicola Salvati.

17 recordsLinked to original sources

A Partial Fay-Herriot model for Small Area Estimation: Estimating district-level consumption in Mozambique

This paper proposes a new small area estimation approach that integrates Partial Least Squares within the Fay-Herriot model to address the challenges posed by high-dimensional and highly correlated auxiliary variables sets. The resulting Partial Fay-Herriot (PFH) estimator constructs supervised components that maximize their association with the target variable, enhancing model stability and predictive efficiency. Monte Carlo simulations demonstrate that PFH estimator achieves lower mean squared error than the standard Fay-Herriot estimator and outperforms principal components-based alternatives while relying on fewer latent dimensions. The methodology is applied to the estimation of district-level per capita consumption in Mozambique, where the survey data source is complemented by large set of correlated census variables. The resulting estimates highlight pronounced geographic heterogeneity and uncover spatial clusters of deprivation. Overall, the findings show that the proposed supervised dimension-reduction approach represents an effective and easily interpretable tool for producing reliable indicators in high-dimensional contexts.

stat.ME

Empirical best prediction of poverty indicators via nested error regression with high dimensional parameters

The Nested Error Regression Model with High-Dimensional Parameters (NERHDP) is extended to address challenges in small area poverty estimation. A robust and flexible framework is proposed to derive empirical best predictors (EBPs) of small area poverty indicators while accommodating heterogeneity in regression coefficients and sampling variances across areas. To mitigate the computational limitations of the existing algorithm, an efficient estimation procedure is introduced, substantially reducing computation time and enhancing scalability for large datasets. A novel approach for generating area-specific poverty estimates in out-of-sample areas is also developed, improving the reliability of synthetic estimates. Uncertainty is quantified through a parametric bootstrap method specifically tailored to the extended model. Under heterogeneous data-generating scenarios, the proposed method yields lower relative bias and relative root mean squared prediction error than existing approaches. The methodology is further illustrated using data from the 2002 Albania Living Standards Measurement Survey, combined with auxiliary information from the 2001 census, to estimate poverty indicators for 374 municipalities.

stat.ME

lqmix: an R package for longitudinal data analysis via linear quantile mixtures

The analysis of longitudinal data gives the chance to observe how unit behaviors change over time, but it also poses a series of issues. These have been the focus of an extensive literature in the context of linear and generalized linear regression moving also, in the last ten years or so, to the context of linear quantile regression for continuous responses. In this paper, we present \texttt{lqmix}, a novel \texttt{R} package that assists in estimating a class of linear quantile regression models for longitudinal data, in the presence of time-constant and/or time-varying, unit-specific, random coefficients, with unspecified distribution. Model parameters are estimated in a maximum likelihood framework via an extended EM algorithm, while parameters' standard errors are derived via a block-bootstrap procedure. The analysis of a benchmark dataset is used to give details on the package functions.

stat.CO

The impact of job stability on monetary poverty in Italy: causal small area estimation

Job stability - encompassing secure contracts, adequate wages, social benefits, and career opportunities - is a critical determinant in reducing monetary poverty, as it provides households with reliable income and enhances economic well-being. This study leverages EU-SILC survey and census data to estimate the causal effect of job stability on monetary poverty across Italian provinces, quantifying its influence and analyzing regional disparities. We introduce a novel causal small area estimation (CSAE) framework that integrates global and local estimation strategies for heterogeneous treatment effect estimation, effectively addressing data sparsity at the provincial level. Furthermore, we develop a general bootstrap scheme to construct reliable confidence intervals, applicable regardless of the method used for estimating nuisance parameters. Extensive simulation studies demonstrate that our proposed estimators outperform classical causal inference methods in terms of stability while maintaining computational scalability for large datasets. Applying this methodology to real-world data, we uncover significant relationships between job stability and poverty across six Italian regions, offering critical insights into regional disparities and their implications for evidence-based policy design.

stat.AP

Effects of model misspecification on small area estimators

Nested error regression models are commonly used to incorporate observational unit specific auxiliary variables to improve small area estimates. When the mean structure of this model is misspecified, there is generally an increase in the mean square prediction error (MSPE) of Empirical Best Linear Unbiased Predictors (EBLUP). Observed Best Prediction (OBP) method has been proposed with the intent to improve on the MSPE over EBLUP. We conduct a Monte Carlo simulation experiment to understand the effect of mispsecification of mean structures on different small area estimators. Our simulation results lead to an unexpected result that OBP may perform very poorly when observational unit level auxiliary variables are used and that OBP can be improved significantly when population means of those auxiliary variables (area level auxiliary variables) are used in the nested error regression model or when a corresponding area level model is used. Our simulation also indicates that the MSPE of OBP in an increasing function of the difference between the sample and population means of the auxiliary variables.

stat.ME

Temporal M-quantile models and robust bias-corrected small area predictors

In small area estimation, it is a smart strategy to rely on data measured over time. However, linear mixed models struggle to properly capture time dependencies when the number of lags is large. Given the lack of published studies addressing robust prediction in small areas using time-dependent data, this research seeks to extend M-quantile models to this field. Indeed, our methodology successfully addresses this challenge and offers flexibility to the widely imposed assumption of unit-level independence. Under the new model, robust bias-corrected predictors for small area linear indicators are derived. Additionally, the optimal selection of the robustness parameter for bias correction is explored, contributing theoretically to the field and enhancing outlier detection. For the estimation of the mean squared error (MSE), a first-order approximation and analytical estimators are obtained under general conditions. Several simulation experiments are conducted to evaluate the performance of the fitting algorithm, the new predictors, and the resulting MSE estimators, as well as the optimal selection of the robustness parameter. Finally, an application to the Spanish Living Conditions Survey data illustrates the usefulness of the proposed predictors.

stat.ME

Accounting for Mismatch Error in Small Area Estimation with Linked Data

In small area estimation different data sources are integrated in order to produce reliable estimates of target parameters (e.g., a mean or a proportion) for a collection of small subsets (areas) of a finite population. Regression models such as the linear mixed effects model or M-quantile regression are often used to improve the precision of survey sample estimates by leveraging auxiliary information for which means or totals are known at the area level. In many applications, the unit-level linkage of records from different sources is probabilistic and potentially error-prone. In this paper, we present adjustments of the small area predictors that are based on either the linear mixed effects model or M-quantile regression to account for the presence of linkage error. These adjustments are developed from a two-component mixture model that hinges on the assumption of independence of the target and auxiliary variable given incorrect linkage. Estimation and inference is based on composite likelihoods and machinery revolving around the Expectation-Maximization Algorithm. For each of the two regression methods, we propose modified small area predictors and approximations for their mean squared errors. The empirical performance of the proposed approaches is studied in both design-based and model-based simulations that include comparisons to a variety of baselines.

stat.ME

Estimating causal quantile exposure response functions via matching

We develop new matching estimators for estimating causal quantile exposure-response functions and quantile exposure effects with continuous treatments. We provide identification results for the parameters of interest and establish the asymptotic properties of the derived estimators. We introduce a two-step estimation procedure. In the first step, we construct a matched data set via generalized propensity score matching, adjusting for measured confounding. In the second step, we fit a kernel quantile regression to the matched set. We also derive a consistent estimator of the variance of the matching estimators. Using simulation studies, we compare the introduced approach with existing alternatives in various settings. We apply the proposed method to Medicare claims data for the period 2012-2014, and we estimate the causal effect of exposure to PM$_{2.5}$ on the length of hospital stay for each zip code of the contiguous United States.

stat.ME

Unified unconditional regression for multivariate quantiles, M-quantiles and expectiles

In this paper, we develop a unified regression approach to model unconditional quantiles, M-quantiles and expectiles of multivariate dependent variables exploiting the multidimensional Huber's function. To assess the impact of changes in the covariates across the entire unconditional distribution of the responses, we extend the work of Firpo et al. (2009) by running a mean regression of the recentered influence function on the explanatory variables. We discuss the estimation procedure and establish the asymptotic properties of the derived estimators. A data-driven procedure is also presented to select the tuning constant of the Huber's function. The validity of the proposed methodology is explored with simulation studies and through an application using the Survey of Household Income and Wealth 2016 conducted by the Bank of Italy.

stat.ME

A Nested Error Regression Model with High Dimensional Parameter for Small Area Estimation

In this paper we propose a flexible nested error regression small area model with high dimensional parameter that incorporates heterogeneity in regression coefficients and variance components. We develop a new robust small area specific estimating equations method that allows appropriate pooling of a large number of areas in estimating small area specific model parameters. We propose a parametric bootstrap and jackknife method to estimate not only the mean squared errors but also other commonly used uncertainty measures such as standard errors and coefficients of variation. We conduct both modelbased and design-based simulation experiments and real-life data analysis to evaluate the proposed methodology

stat.ME

Causal Inferences in Small Area Estimation

When doing impact evaluation and making causal inferences, it is important to acknowledge the heterogeneity of the treatment effects for different domains (geographic, socio-demographic, or socio-economic). If the domain of interest is small with regards to its sample size (or even zero in some cases), then the evaluator has entered the small area estimation (SAE) dilemma. Based on the modification of the Inverse Propensity Weighting estimator and the traditional small area predictors, the paper proposes a new methodology to estimate area specific average treatment effects for unplanned domains. By means of these methods we can also provide a map of policy impacts, that can help to better target the treatment group(s). We develop analytical Mean Squared Error (MSE) estimators of the proposed predictors. An extensive simulation analysis, also based on real data, shows that the proposed techniques in most cases lead to more efficient estimators.

stat.ME

Scale estimation and data-driven tuning constant selection for M-quantile regression

M-quantile regression is a general form of quantile-like regression which usually utilises the Huber influence function and corresponding tuning constant. Estimation requires a nuisance scale parameter to ensure the M-quantile estimates are scale invariant, with several scale estimators having previously been proposed. In this paper we assess these scale estimators and evaluate their suitability, as well as proposing a new scale estimator based on the method of moments. Further, we present two approaches for estimating data-driven tuning constant selection for M-quantile regression. The tuning constants are obtained by i) minimising the estimated asymptotic variance of the regression parameters and ii) utilising an inverse M-quantile function to reduce the effect of outlying observations. We investigate whether data-driven tuning constants, as opposed to the usual fixed constant, for instance, at c=1.345, can improve the efficiency of the estimators of M-quantile regression parameters. The performance of the data-driven tuning constant is investigated in different scenarios using model-based simulations. Finally, we illustrate the proposed methods using a European Union Statistics on Income and Living Conditions data set.

stat.ME

Small Area Estimation with Linked Data

In Small Area Estimation data linkage can be used to combine values of the variableof interest from a national survey with values of auxiliary variables obtained from another source like a population register. Linkage errors can induce bias when fitting regression models; moreover, they can create non-representative outliers in the linked data in addition to the presence of potential representative outliers. In this paper we adopt a secondary analyst's point view, assuming limited information is available on the linkage process, and we develop small area estimators based on linear mixed and linear M-quantile models to accommodate linked data containing a mix of both types of outliers. We illustrate the properties of these small area estimators, as well as estimators of their mean squared error, by means of model-based and design-based simulation experiments. These experiments show that the proposed predictors can lead to more efficient estimators when there is linkage error. Furthermore, the proposed mean-squared error estimation methods appear to perform well.

stat.ME

Semi-Parametric Empirical Best Prediction for small area estimation of unemployment indicators

The Italian National Institute for Statistics regularly provides estimates of unemployment indicators using data from the Labor Force Survey. However, direct estimates of unemployment incidence cannot be released for Local Labor Market Areas. These are unplanned domains defined as clusters of municipalities; many are out-of-sample areas and the majority is characterized by a small sample size, which render direct estimates inadequate. The Empirical Best Predictor represents an appropriate, model-based, alternative. However, for non-Gaussian responses, its computation and the computation of the analytic approximation to its Mean Squared Error require the solution of (possibly) multiple integrals that, generally, have not a closed form. To solve the issue, Monte Carlo methods and parametric bootstrap are common choices, even though the computational burden is a non trivial task. In this paper, we propose a Semi-Parametric Empirical Best Predictor for a (possibly) non-linear mixed effect model by leaving the distribution of the area-specific random effects unspecified and estimating it from the observed data. This approach is known to lead to a discrete mixing distribution which helps avoid unverifiable parametric assumptions and heavy integral approximations. We also derive a second-order, bias-corrected, analytic approximation to the corresponding Mean Squared Error. Finite sample properties of the proposed approach are tested via a large scale simulation study. Furthermore, the proposal is applied to unit-level data from the 2012 Italian Labor Force Survey to estimate unemployment incidence for 611 Local Labor Market Areas using auxiliary information from administrative registers and the 2011 Census.

stat.ME

The use of sampling weights in the M-quantile random-effects regression: an application to PISA mathematics scores

M-quantile random-effects regression represents an interesting approach for modelling multilevel data when the interest of researchers is focused on the conditional quantiles. When data are based on complex survey designs, sampling weights have to be incorporate in the analysis. A pseudo-likelihood approach for accommodating sampling weights in the M-quantile random-effects regression is presented. The proposed methodology is applied to the Italian sample of the "Program for International Student Assessment 2015" survey in order to study the gender gap in mathematics at various quantiles of the conditional distribution. Findings offer a possible explanation of the low share of females in "Science, Technology, Engineering and Mathematics" sectors.

math.ST

M-quantile regression for multivariate longitudinal data: analysis of the Millennium Cohort Study data

We propose a M-quantile regression model for the analysis of multivariate, continuous, longitudinal data. M-quantile regression represents an appealing alternative to standard regression models, as it combines the robustness of quantile and the efficiency of expectile regression, providing a complete picture of the response variable distribution. Discrete, individual-specific, random parameters are used to account for both dependence within the same response recorded at different times and association between different responses observed on the same sample unit at a given time. A suitable parametrisation is also introduced in the linear predictor to account for possible dependence between the individual specific random parameters and the vector of observed covariates, that is to account for endogeneity of some covariates. An extended EM algorithm is proposed to derive model parameter estimates under a maximum likelihood approach. The model is applied to the analysis of the strengths and difficulties questionnaire scores from the Millennium Cohort Study in the UK.

stat.ME

Disease Mapping via Negative Binomial Regression M-quantiles

We introduce a semi-parametric approach to ecological regression for disease mapping, based on modelling the regression M-quantiles of a Negative Binomial variable. The proposed method is robust to outliers in the model covariates, including those due to measurement error, and can account for both spatial heterogeneity and spatial clustering. A simulation experiment based on the well-known Scottish lip cancer data set is used to compare the M-quantile modelling approach and a random effects modelling approach for disease mapping. This suggests that the M-quantile approach leads to predicted relative risks with smaller root mean square error than standard disease mapping methods. The paper concludes with an illustrative application of the M-quantile approach, mapping low birth weight incidence data for English Local Authority Districts for the years 2005-2010.

stat.ME