SearcharxivSearch

arXiv subjects

Summer Wang

Publications and source records attributed to Summer Wang.

2 recordsLinked to original sources

Estimating Negative Income Distributions via Data Fusion with Vine Copula-based Imputation

Imputation-based data fusion combines datasets by imputing missing variables in one dataset using information from the other. The validity of imputation-based data fusion relies on accurately preserving both the imputed distributions and the underlying multivariate dependence structure across data sources. However, traditional imputation-based data fusion methods often struggle to preserve the dependence structure, particularly when marginal distributions differ or when complex, nonlinear multivariate dependencies exist. Inadequate preservation of these dependencies can result in distorted joint distributions and biased inference in the fused data. To address this challenge, this study introduces a novel imputation-based data fusion framework that utilises conditional sampling from C-vine and D-vine copulas as the imputation mechanism. This approach flexibly models pairwise and higher-order dependencies while accommodating heterogeneous marginal distributions, thereby generating imputations that more accurately maintain the distribution of the multivariate data. The proposed method is validated through a simulation study using survey data and a real-data application that estimates the negative income distribution in survey data using information from administrative tax records. Results indicate that the proposed method outperforms existing approaches in preserving both the imputed distributions and the dependence structure across datasets.

stat.ME

A note on the optimum allocation of resources to follow up unit nonrespondents in probability

Common practice to address nonresponse in probability surveys in National Statistical Offices is to follow up every nonrespondent with a view to lifting response rates. As response rate is an insufficient indicator of data quality, it is argued that one should follow up nonrespondents with a view to reducing the mean squared error (MSE) of the estimator of the variable of interest. In this paper, we propose a method to allocate the nonresponse follow-up resources in such a way as to minimise the MSE under a quasi-randomisation framework. An example to illustrate the method using the 2018/19 Rural Environment and Agricultural Commodities Survey from the Australian Bureau of Statistics is provided.

stat.ME