Searcharxiv⌕ Search

arXiv subjects

Guillaume Chauvet

Publications and source records attributed to Guillaume Chauvet.

17 recordsLinked to original sources

Inference for two-stage sampling in spatial surveys

This paper develops a design-based asymptotic theory for two-stage sampling over a continuous spatial domain. The target parameters are integral totals, or smooth functions of such totals, defined over a fixed bounded territory partitioned into an increasingly fine collection of primary sampling units. Within this fixed-area framework, we derive the order of the variance components of the Horvitz-Thompson estimator when secondary sampling units are points selected from continuous sub-regions. We establish design consistency of the estimator, as well as consistency of variance estimators under explicit regularity conditions on inclusion probabilities, sampling densities, and the study variable. The results are extended to plug-in estimators of smooth scaleinvariant functions of totals. For high-entropy first-stage designs, we further show that a H{á}jek-type variance estimator based only on first-order inclusion probabilities is consistent, and we prove asymptotic normality of both total and plug-in estimators. A simulation study inspired by forest inventory applications illustrates the finite-sample performance of the proposed estimators and variance estimators.

math.ST↗

Reflection on modern methods: a note on variance estimation when using inverse probability weighting to handle attrition in cohort studies

The inverse probability weighting (IPW) method is used to handle attrition in association analyses derived from cohort studies. It consists in weighting the respondents at a given follow-up by their inverse probability to participate. Weights are estimated first and then used in a weighted association model. When the IPW method is used, instead of using a so-called na{ï}ve variance estimator, the literature recommends using a robust variance estimator. However, the latter may overestimate the variance because the weights are considered known rather than estimated. In this note, we develop, by a linearization technique, an estimator accounting for the weight estimation phase and explain how it differs from na{ï}ve and robust variance estimators. We compare the three variance estimators through simulations under several MAR and MNAR scenarios. We found that both the robust and linearized variance estimators were approximately unbiased, even in MNAR scenarios. The naive variance estimator severely underestimated the variance. We encourage researchers to be careful with variance estimation when using the IPW method, avoiding na{ï}ve estimator and opting for a robust or linearized estimator. R and SAS codes are provided to implement them in their own studies.

stat.AP↗

A cautionary note on the Hanurav-Vijayan sampling algorithm

We consider the Hanurav-Vijayan sampling design, which is the default method programmed in the SURVEYSELECT procedure of the SAS software. We prove that it is equivalent to the Sunter procedure, but is capable of handling any set of inclusion probabilities. We prove that the Horvitz-Thompson estimator is not generally consistent under this sampling design. We propose a conditional Horvitz-Thompson estimator, and prove its consistency under a non-standard assumption on the first-order inclusion probabilities. Since this assumption seems difficult to control in practice, we recommend not to use the Hanurav-Vijayan sampling design.

stat.ME↗

Exponential inequalities for sampling designs

In this work we introduce a general approach, based on the mar-tingale representation of a sampling design and Azuma-Hoeffding's inequality , to derive exponential inequalities for the difference between a Horvitz-Thompson estimator and its expectation. Applying this idea, we establish such inequalities for Chao's procedure, Till{é}'s elimination procedure, the generalized Midzuno method as well as for Brewer's method. As a by-product, we prove that the first three sampling designs are (conditionally) negatively associated. For such sampling designs, we show that that the inequality we obtain is usually sharper than the one obtained by applying known results for negatively associated random variables.

math.ST↗

Closed-form variance estimators for weighted and stratified dose-response function estimators using generalized propensity score

Propensity score methods are widely used in observational studies for evaluating marginal treatment effects. The generalized propensity score (GPS) is an extension of the propensity score framework, historically developed in the case of binary exposures, for use with quantitative or continuous exposures. In this paper, we proposed variance esti-mators for treatment effect estimators on continuous outcomes. Dose-response functions (DRF) were estimated through weighting on the inverse of the GPS, or using stratification. Variance estimators were evaluated using Monte Carlo simulations. Despite the use of stabilized weights, the variability of the weighted estimator of the DRF was particularly high, and none of the variance estimators (a bootstrap-based estimator, a closed-form estimator especially developped to take into account the estimation step of the GPS, and a sandwich estimator) were able to adequately capture this variability, resulting in coverages below to the nominal value, particularly when the proportion of the variation in the quantitative exposure explained by the covariates was 1 large. The stratified estimator was more stable, and variance estima-tors (a bootstrap-based estimator, a pooled linearized estimator, and a pooled model-based estimator) more efficient at capturing the empirical variability of the parameters of the DRF. The pooled variance estimators tended to overestimate the variance, whereas the bootstrap estimator, which intrinsically takes into account the estimation step of the GPS, resulted in correct variance estimations and coverage rates. These methods were applied to a real data set with the aim of assessing the effect of maternal body mass index on newborn birth weight.

stat.AP↗

Properties of Chromy's sampling procedure

Chromy (1979) proposed a unequal probability sampling algorithm, which enables to select a sample in one pass of the sampling frame only. This is the default sequential method used in the SURVEYSELECT procedure of the SAS software. In this article, we study the properties of Chromy sampling. We prove that the Horvitz-Thompson is asymptotically normally distributed, and give an explicit expression for the second-order inclusion probabilities. This makes it possible to estimate the variance unbiasedly for the randomized version of the method programmed in the SURVEYSELECT procedure.

math.ST↗

Preserving the distribution function in surveys in case of imputation for zero inflated data

Item non-response in surveys is usually handled by single imputation, whose main objective is to reduce the non-response bias. Imputation methods need to be adapted to the study variable. For instance, in business surveys, the interest variables often contain a large number of zeros. Motivated by a mixture regression model, we propose two imputation procedures for such data and study their statistical properties. We show that these procedures preserve the distribution function if the imputation model is well specified. The results of a simulation study illustrate the good performance of the proposed methods in terms of bias and mean square error.

stat.ME↗

Inference for two-stage sampling designs with application to a panel for urban policy

Two-stage sampling designs are commonly used for household and health surveys. To produce reliable estimators with assorted confidence intervals, some basic statistical properties like consistency and asymptotic normality of the Horvitz-Thompson estimator are desirable, along with the consistency of assorted variance estimators. These properties have been mainly studied for single-stage sampling designs. In this work, we prove the consistency of the Horvitz-Thompson estimator and of associated variance estimators for a general class of two-stage sampling designs, under mild assumptions. We also study two-stage sampling with a large entropy sampling design at the first stage, and prove that the Horvitz-Thompson estimator is asymptotically normally distributed through a coupling argument. When the first-stage sampling fraction is negligible, simplified variance estimators which do not require estimating the variance within the Primary Sampling Units are proposed, and shown to be consistent. An application to a panel for urban policy, which is the initial motivation for this work, is also presented.

stat.ME↗

Exact balanced random imputation for sample survey data

Surveys usually suffer from non-response, which decreases the effective sample size. Item non-response is typically handled by means of some form of random imputation if we wish to preserve the distribution of the imputed variable. This leads to an increased variability due to the imputation variance, and several approaches have been proposed for reducing this variability. Balanced imputation consists in selecting residuals at random at the imputation stage, in such a way that the imputation variance of the estimated total is eliminated or at least significantly reduced. In this work, we propose an implementation of balanced random imputation which enables to fully eliminate the imputation variance. Following the approach in Cardot et al. (2013), we consider a regularized imputed estimator of a total and of a distribution function, and we prove that they are consistent under the proposed imputation method. Some simulation results support our findings.

stat.ME↗

Coupling methods for multistage sampling

Multistage sampling is commonly used for household surveys when there exists no sampling frame, or when the population is scattered over a wide area. Multistage sampling usually introduces a complex dependence in the selection of the final units, which makes asymptotic results quite difficult to prove. In this work, we consider multistage sampling with simple random without replacement sampling at the first stage, and with an arbitrary sampling design for further stages. We consider coupling methods to link this sampling design to sampling designs where the primary sampling units are selected independently. We first generalize a method introduced by [Magyar Tud. Akad. Mat. Kutató Int. Közl. 5 (1960) 361-374] to get a coupling with multistage sampling and Bernoulli sampling at the first stage, which leads to a central limit theorem for the Horvitz--Thompson estimator. We then introduce a new coupling method with multistage sampling and simple random with replacement sampling at the first stage. When the first-stage sampling fraction tends to zero, this method is used to prove consistency of a with-replacement bootstrap for simple random without replacement sampling at the first stage, and consistency of bootstrap variance estimators for smooth functions of totals.

math.ST↗

Joint imputation procedures for categorical variables

Marginal imputation, which consists of imputing each item requiring imputation separately, is often used in surveys. This type of imputation procedures leads to asymptotically unbiased estimators of simple parameters such as population totals (or means), but tends to distort relationships between variables. As a result, it generally leads to biased estimators of bivariate parameters such as coefficients of correlation or odd-ratios. Household and social surveys typically collect categorical variables, for which missing values are usually handled by nearest-neighbour imputation or random hot-deck imputation. In this paper, we propose a simple random imputation procedure, closely related to random hot-deck imputation, which succeeds in preserving the relationship between categorical variables. Also, a fully efficient version of the latter procedure is proposed. A limited simulation study compares several estimation procedures in terms of relative bias and relative efficiency.

stat.ME↗

Estimation under cross-classified sampling with application to a childhood survey

The cross-classified sampling design consists in drawing samples from a two-dimension population, independently in each dimension. Such design is commonly used in consumer price index surveys and has been recently applied to draw a sample of babies in the French ELFE survey, by crossing a sample of maternity units and a sample of days. We propose to derive a general theory of estimation for this sampling design. We consider the Horvitz-Thompson estimator for a total, and show that the cross-classified design will usually result in a loss of efficiency as compared to the widespread two-stage design. We obtain the asymptotic distribution of the Horvitz-Thompson estimator, and several unbiased variance estimators. Facing the problem of possibly negative values, we propose simplified non-negative variance estimators and study their bias under a super-population model. The proposed estimators are compared for totals and ratios on simulated data. An application on real data from the ELFE survey is also presented, and we make some recommendations. Supplementary materials are available online.

math.ST↗

Martingale central-limit theorems for pivotal sampling

Ordered pivotal sampling is one of the simplest algorithm to perform without-replacement unequal probability sampling. It has found uses in the context of longitudinal surveys and spatial sampling, and enables in particular a good spatial balance of the selected units. In this work, we follow the approach proposed by Ohlsson~(1986), and apply a martingale central-limit theorem to prove the asymptotic normality of the Horvitz-Thompson estimator under a design-based approach, and under a model-assisted approach. In particular, our model assumptions allow for correlations between values, which is of particular interest for applications in spatial sampling.

math.ST↗

A note on the consistency of the Narain-Horvitz-Thompson estimator

For the Narain-Horvitz-Thompson estimator to have usual asymptotic properties such as consistency, some conditions on the sampling design and on the variable of interest are needed. Cardot et al. (2010) give some sufficient conditions for the mean square consistency, but one of them is usually difficult to prove or does not hold for some unequal probability sampling designs. We propose alternative conditions for the mean square consistency of the Narain-Horvitz-Thompson estimator. A specific result is also proved in case when a martingale sampling algorithm is used, which implies consistency under a fast algorithm for the cube method.

stat.ME↗

On a characterization of ordered pivotal sampling

When auxiliary information is available at the design stage, samples may be selected by means of balanced sampling. Deville and Tille proposed in 2004 a general algorithm to perform balanced sampling, named the cube method. In this paper, we are interested in a particular case of the cube method named pivotal sampling, and first described by Deville and Tille in 1998. We show that this sampling algorithm, when applied to units ranked in a fixed order, is equivalent to Deville's systematic sampling, in the sense that both algorithms lead to the same sampling design. This characterization enables the computation of the second-order inclusion probabilities for pivotal sampling. We show that the pivotal sampling enables to take account of an appropriate ordering of the units to achieve a variance reduction, while limiting the loss of efficiency if the ordering is not appropriate.

math.ST↗