SearcharxivSearch

arXiv subjects

David Haziza

Publications and source records attributed to David Haziza.

8 recordsLinked to original sources

Agnostic Model-Assisted Estimation with Machine Learning for Survey Data

Model-assisted estimation uses prediction rules to improve the efficiency of estimators of finite population parameters while retaining design-based inference. Although flexible prediction methods have been considered, existing theoretical results are largely method-specific. We develop a learner-agnostic framework that replaces separate analyses for individual learners with general conditions on the sampling design and prediction error. We connect design-aware and design-agnostic cross-fitting and characterize the sampling designs under which they yield conditional independence across folds. Under suitable conditions, conditional weighting gives exact design-unbiasedness. We establish first-order equivalence to oracle estimators, leading to design consistency and asymptotic normality, and clarify when conditional and original inclusion probabilities yield the same first-order behavior. We propose consistent variance estimators based on cross-fitted residuals and construct asymptotically valid confidence intervals. Under additional model and regularity conditions, we establish asymptotic optimality through attainment of the Godambe--Joshi lower bound. Simulations show that cross-fitting substantially reduces finite-sample bias and improves variance estimation and coverage with adaptive learners.

stat.ME

Machine learning methods for finite population parameter estimation in survey sampling

This pedagogical review examines the use of machine learning methods in finite-population inference for survey sampling, with an emphasis on design-based validity and statistical inference. While flexible prediction tools offer substantial gains in estimation accuracy, they also introduce important challenges, primarily due to the dependence between the fitted predictors and the sample. We focus on settings in which such predictions enter survey estimation through model-assisted estimation, item nonresponse imputation, and unit nonresponse adjustment. For model-assisted estimation and item nonresponse, we show how cross-fitting and Neyman-orthogonal estimating equations can adapt ideas from double/debiased machine learning to survey data, allowing the use of high-dimensional or nonparametric learners while preserving root-n consistency and asymptotic normality under suitable conditions. In contrast, for unit nonresponse, standard inverse-probability weighting remains outcome-agnostic and operationally attractive, but this same feature makes doubly robust and orthogonal constructions harder to deploy in official statistics. We also briefly discuss related developments in small area estimation and probability/nonprobability data integration. Overall, the paper highlights both the promise of machine learning and the fundamental inferential challenges it raises for survey practice.

stat.ME

Variable Selection for Linear Regression Imputation in Surveys

Survey sampling is concerned with the estimation of finite population parameters. In practice, survey data suffer from item nonresponse, which is commonly handled through imputation, i.e., replacing missing values with predicted values. As a result, the properties of the resulting imputed estimator depend critically on the properties of the prediction method used. In turn, prediction methods themselves depend on the choice of variables and tuning parameters used to fit the imputation model. In this article, we study the problem of variable selection for linear regression imputation. Although variable selection has been widely studied across many fields, primarily for identification or prediction, its role in imputation for survey data has received comparatively little attention. We introduce the notion of an optimal imputation model defined through an oracle loss function and show that, with probability tending to one, the optimal model coincides with the true model. We also examine the consequences of using misspecified models -- either omitting relevant covariates or including irrelevant ones -- on consistency and asymptotic variance. We then develop a complete methodological framework for constructing confidence intervals after model selection. The proposed confidence intervals are shown to be asymptotically valid and optimal among all candidate models. Simulation studies indicate that the proposed methodology performs well in finite samples.

stat.ME

Model-assisted estimation through random forests in finite population sampling

In surveys, the interest lies in estimating finite population parameters such as population totals and means. In most surveys, some auxiliary information is available at the estimation stage. This information may be incorporated in the estimation procedures to increase their precision. In this article, we use random forests to estimate the functional relationship between the survey variable and the auxiliary variables. In recent years, random forests have become attractive as National Statistical Offices have now access to a variety of data sources, potentially exhibiting a large number of observations on a large number of variables. We establish the theoretical properties of model-assisted procedures based on random forests and derive corresponding variance estimators. A model-calibration procedure for handling multiple survey variables is also discussed. The results of a simulation study suggest that the proposed point and estimation procedures perform well in term of bias, efficiency, and coverage of normal-based confidence intervals, in a wide variety of settings. Finally, we apply the proposed methods using data on radio audiences collected by Médiamétrie, a French audience company.

stat.ME

Model-assisted estimation in high-dimensional settings for survey data

Model-assisted estimators have attracted a lot of attention in the last three decades. These estimators attempt to make an efficient use of auxiliary information available at the estimation stage. A working model linking the survey variable to the auxiliary variables is specified and fitted on the sample data to obtain a set of predictions, which are then incorporated in the estimation procedures. A nice feature of model-assisted procedures is that they maintain important design properties such as consistency and asymptotic unbiasedness irrespective of whether or not the working model is correctly specified. In this article, we examine several model-assisted estimators from a design-based point of view and in a high-dimensional setting, including penalized estimators and tree-based estimators. We conduct an extensive simulation study using data from the Irish Commission for Energy Regulation Smart Metering Project, in order to assess the performance of several model-assisted estimators in terms of bias and efficiency in this high-dimensional data set.

stat.ME

Imputation procedures in surveys using nonparametric and machine learning methods: an empirical comparison

Nonparametric and machine learning methods are flexible methods for obtaining accurate predictions. Nowadays, data sets with a large number of predictors and complex structures are fairly common. In the presence of item nonresponse, nonparametric and machine learning procedures may thus provide a useful alternative to traditional imputation procedures for deriving a set of imputed values. In this paper, we conduct an extensive empirical investigation that compares a number of imputation procedures in terms of bias and efficiency in a wide variety of settings, including high-dimensional data sets. The results suggest that a number of machine learning procedures perform very well in terms of bias and efficiency.

stat.ME

Efficient multiply robust imputation in the presence of influential units in surveys

Item nonresponse is a common issue in surveys. Because unadjusted estimators may be biased in the presence of nonresponse, it is common practice to impute the missing values with the objective of reducing the nonresponse bias as much as possible. However, commonly used imputation procedures may lead to unstable estimators of population totals/means when influential units are present in the set of respondents. In this article, we consider the class of multiply robust imputation procedures that provide some protection against the failure of underlying model assumptions. We develop an efficient version of multiply robust estimators based on the concept of conditional bias, a measure of influence. We present the results of a simulation study to show the benefits of the proposed method in terms of bias and efficiency.

stat.ME

Joint imputation procedures for categorical variables

Marginal imputation, which consists of imputing each item requiring imputation separately, is often used in surveys. This type of imputation procedures leads to asymptotically unbiased estimators of simple parameters such as population totals (or means), but tends to distort relationships between variables. As a result, it generally leads to biased estimators of bivariate parameters such as coefficients of correlation or odd-ratios. Household and social surveys typically collect categorical variables, for which missing values are usually handled by nearest-neighbour imputation or random hot-deck imputation. In this paper, we propose a simple random imputation procedure, closely related to random hot-deck imputation, which succeeds in preserving the relationship between categorical variables. Also, a fully efficient version of the latter procedure is proposed. A limited simulation study compares several estimation procedures in terms of relative bias and relative efficiency.

stat.ME