SearcharxivSearch

arXiv subjects

Caren Hasler

Publications and source records attributed to Caren Hasler.

7 recordsLinked to original sources

In Search of the Most Balanced Sampling Design

Balanced sampling aims to select random samples in which the estimated totals of the auxiliary variables, weighted by the inverse of the inclusion probabilities, correspond as closely as possible to the known population totals. While several methods, such as rejective sampling, rerandomization, and the cube method, have been proposed to improve balance, identifying the most balanced sampling design under fixed inclusion probabilities remains a challenging combinatorial problem. This problem can be formulated as a linear program defined over the set of all possible samples, but the number of samples grows exponentially with population size, making exact optimization infeasible except for very small populations. To address this issue, we propose a heuristic approach based on a genetic algorithm that iteratively improves the balance of sampling designs by combining minimum support designs with highly balanced candidate samples. Although optimality cannot be guaranteed, the proposed method can substantially improve balance relative to standard procedures such as the cube method. The approach is applicable to both survey sampling and experimental design.

stat.ME

Consistency of the $k$-Nearest Neighbor Regressor under Complex Survey Designs

We study the consistency of the $k$-nearest neighbor regressor under complex survey designs. While consistency results for this algorithm are well established for independent and identically distributed data, corresponding results for complex survey data are lacking. We show that the $k$-nearest neighbor regressor is consistent under regularity conditions on the sampling design and the distribution of the data. We derive lower bounds for the rate of convergence and show that these bounds exhibit the curse of dimensionality, as in the independent and identically distributed setting. Empirical studies based on simulated and real data illustrate our theoretical findings.

stat.ML

Calibration with Bagging of the Principal Components on a Large Number of Auxiliary Variables

Calibration is a widely used method in survey sampling to adjust weights so that estimated totals of some chosen calibration variables match known population totals or totals obtained from other sources. When a large number of auxiliary variables are included as calibration variables, the variance of the total estimator can increase, and the calibration weights can become highly dispersed. To address these issues, we propose a solution inspired by bagging and principal component decomposition. With our approach, the principal components of the auxiliary variables are constructed. Several samples of calibration variables are selected without replacement and with unequal probabilities from among the principal components. For each sample, a system of weights is obtained. The final weights are the average weights of these different weighting systems. With our proposed method, it is possible to calibrate exactly for some of the main auxiliary variables. For the other auxiliary variables, the weights cannot be calibrated exactly. The proposed method allows us to obtain a total estimator whose variance does not explode when new auxiliary variables are added and to obtain very low scatter weights. Finally, our proposed method allows us to obtain a single weighting system that can be applied to several variables of interest of a survey.

stat.ME

Quasi Model-Assisted Estimators under Nonresponse in Sample Surveys

In the presence of auxiliary information, model-assisted estimators rely on a working model linking the variable of interest to the auxiliary variables in order to improve the efficiency of the Horvitz-Thompson estimator. Model-assisted estimators cannot be directly computed with nonresponse since the values of the variable of interest is missing for a part of the sample units. In this article, we present and study a class of quasi-model-assisted estimators that extend model-assisted estimators to settings with non-ignorable nonresponse. These estimators combine a working model and a response model. The former is used to improve the efficiency, the latter to reweight the nonrespondents. A wide range of statistical learning methods can be used to estimate either of these models. We show that several well-known existing estimators are particular cases of quasi-model-assisted estimators. We examine the behavior of these estimators through a simulation study. The results illustrate how these estimators remain competitive in terms of bias and variance when one of the two models is poorly specified.

stat.ME

Inference from Sampling with Response Probabilities Estimated via Calibration

A solution to control for nonresponse bias consists of multiplying the design weights of respondents by the inverse of estimated response probabilities to compensate for the nonrespondents. Maximum likelihood and calibration are two approaches that can be applied to obtain estimated response probabilities. We consider a common framework in which these approaches can be compared. We develop an asymptotic study of the behavior of the resulting estimator when calibration is applied. A logistic regression model for the response probabilities is postulated. Missing at random and unclustered data are supposed. Three main contributions of this work are: 1) we show that the estimators with the response probabilities estimated via calibration are asymptotically equivalent to unbiased estimators and that a gain in efficiency is obtained when estimating the response probabilities via calibration as compared to the estimator with the true response probabilities, 2) we show that the estimators with the response probabilities estimated via calibration are doubly robust to model misspecification and explain why double robustness is not guaranteed when maximum likelihood is applied, and 3) we discuss and illustrate problems related to response probabilities estimation, namely existence of a solution to the estimating equations, problems of convergence, and extreme weights. We explain and illustrate why the first aforementioned problem is more likely with calibration than with maximum likelihood estimation. We present the results of a simulation study in order to illustrate these elements.

stat.ME

Nonparametric imputation method for nonresponse in surveys

Many imputation methods are based on statistical models that assume that the variable of interest is a noisy observation of a function of the auxiliary variables or covariates. Misspecification of this model may lead to severe errors in estimates and to misleading conclusions. A new imputation method for item nonresponse in surveys is proposed based on a nonparametric estimation of the functional dependence between the variable of interest and the auxiliary variables. We consider the use of smoothing spline estimation within an additive model framework to flexibly build an imputation model in the case of multiple auxiliary variables. The performance of our method is assessed via numerical experiments involving simulated and real data.

stat.ME

Balanced $k$-nearest neighbor imputation

In order to overcome the problem of item nonresponse, random imputation methods are often used because they tend to preserve the distribution of the imputed variable. Among the random imputation methods, the random hot-deck has the interesting property of imputing observed values. A new random hot-deck imputation method is proposed. The key innovation of this method is that the selection of donors is viewed as a sampling problem and uses calibration and balanced sampling. This approach makes it possible to select donors such that if the auxiliary variables were imputed, their estimated totals would not change. As a consequence, very accurate and stable totals estimations can be obtained. Moreover, the method is based on a nonparametric procedure. Donors are selected in neighborhoods of recipients. In this way, the missing value of a recipient is replaced with an observed value of a similar unit. This new approach is very flexible and can greatly improve the quality of estimations. Also, this method is unbiased under very different models and is thus resistant to model misspecification. Finally, the new method makes it possible to introduce edit rules while imputing.

stat.ME