SearcharxivSearch

arXiv subjects

Anders Holmberg

Publications and source records attributed to Anders Holmberg.

6 recordsLinked to original sources

A Total Statistical Error Framework for Comparing Census Data Collection Methods

Population censuses increasingly rely on imputation to assign usual-residence addresses for non-responding dwellings, yet no formal statistical framework has existed for comparing competing imputation methods on their combined coverage and address-accuracy performance. We develop such a framework within a Total Statistical Error (TSE) paradigm that is not restricted to survey-based data collection and applies equally to register-based and administrative data sources. The central quantity is a unit-level binary correctness indicator, equal to one if and only if a person is both enumerated and assigned to the correct usual-residence address; its complement is the unit TSE, and aggregating over the population yields a correctness rate that serves as the basis for head-to-head method comparison. We distinguish two assessment paradigms: Paradigm I, in which a post-enumeration survey (PES) provides reference values for a probability sample, and Paradigm II, in which a complete benchmark makes the correctness rate directly computable. We illustrate the framework using 2021 Australian census microdata, comparing k-nearest-neighbour and random forest imputation under a not-missing-at-random (NMAR) mechanism. A key finding is that dependent nonresponse between the census and a simulated PES causes naive Paradigm I subgroup estimates to be severely biased when subgroup deletion rates are small. Augmenting the PES response propensity model with the census nonresponse indicator reduces this bias by approximately 90%; simulation experiments show that the nonresponse indicator alone drives the correction, with the additionally included imputed set A values contributing negligible further reduction. This is the preferred strategy whenever dependent nonresponse between the census and PES is a concern.

stat.ME

On design-unbiased algorithmic Machine Learning

Machine Learning (ML) algorithms, such as k-Nearest Neighbours (kNN) or random forest, eschew the ideal of true data models in favour of predictive performance. However, minimising the MSE or F-score cannot lead to unbiasedness directly, which is important in many situations such as official statistics. We study the conditions of algorithmic ML, other than the existence and knowledge of true data models, which lead to unbiased prediction or classification for a given finite population, including how the training data may be sampled from the population, how a trained prediction algorithm can be tuned to achieve unbiased prediction or classification for that population, and how the performance of out-of-sample prediction or classification can be assessed unbiasedly. The inference is based on the known probability design of samples and training sets, rather than any assumed distributions or models.

cs.LG

Estimating Propensities of Selection for Big Datasets via Data Integration

Big data presents potential but unresolved value as a source for analysis and inference. However,selection bias, present in many of these datasets, needs to be accounted for so that appropriate inferences can be made on the target population. One way of approaching the selection bias issue is to first estimate the propensity of inclusion in the big dataset for each member of the big dataset, and then to apply these propensities in an inverse probability weighting approach to produce population estimates. In this paper, we provide details of a new variant of existing propensity score estimation methods that takes advantage of the ability to integrate the big data with a probability sample. We compare the ability of this method to produce efficient inferences for the target population with several alternative methods through an empirical study.

stat.ME

An Empirical Comparison of Methods to Produce Business Statistics Using Non-Probability Data

There is a growing trend among statistical agencies to explore non-probability data sources for producing more timely and detailed statistics, while reducing costs and respondent burden. Coverage and measurement error are two issues that may be present in such data. The imperfections may be corrected using available information relating to the population of interest, such as a census or a reference probability sample. In this paper, we compare a wide range of existing methods for producing population estimates using a non-probability dataset through a simulation study based on a realistic business population. The study was conducted to examine the performance of the methods under different missingness and data quality assumptions. The results confirm the ability of the methods examined to address selection bias. When no measurement error is present in the non-probability dataset, a screening dual-frame approach for the probability sample tends to yield lower sample size and mean squared error results. The presence of measurement error and/or nonignorable missingness increases mean squared errors for estimators that depend heavily on the non-probability data. In this case, the best approach tends to be to fall back to a model-assisted estimator based on the probability sample.

stat.ME

A note on the optimum allocation of resources to follow up unit nonrespondents in probability

Common practice to address nonresponse in probability surveys in National Statistical Offices is to follow up every nonrespondent with a view to lifting response rates. As response rate is an insufficient indicator of data quality, it is argued that one should follow up nonrespondents with a view to reducing the mean squared error (MSE) of the estimator of the variable of interest. In this paper, we propose a method to allocate the nonresponse follow-up resources in such a way as to minimise the MSE under a quasi-randomisation framework. An example to illustrate the method using the 2018/19 Rural Environment and Agricultural Commodities Survey from the Australian Bureau of Statistics is provided.

stat.ME

Linking Administrative Data: An Evolutionary Schema

Statistics New Zealand (Stats NZ) has committed unreservedly to an administrative data first policy. Thus, all new methods used at Stats NZ are to be viewed within this context and discussing strategies for using administrative data is an integral part of every working day. As statistical methodologists, the three authors were drawn into these discussions. Like most methodologists, the authors see surveys and the publications of their results as a process where estimation is the key tool to achieve the final goal of an accurate statistical output. Randomness and sampling exists to support this goal, and early on it was clear to us that the incoming it-is-what-it-is data sources were not randomly selected. These sources were obviously biased and thus would produce biased estimates. So, we set out to design a strategy to deal with this issue. This led us to the concept of representativeness which is closely related to statistical bias but has a wider context invoking both randomness and judgement. The representativeness issue was the principal question that we set out to answer. The necessary components that we gathered for our solution are summarized in the paper. Keywords: Representativeness, Timeline Databases, Statistical Registers, Estimation

stat.ME