Searcharxiv⌕ Search

arXiv subjects

Andrius Čiginas

Publications and source records attributed to Andrius Čiginas.

9 recordsLinked to original sources

Diagnostics-guided variance-inflated Fay-Herriot estimation from non-probability samples

Non-probability data sources are increasingly considered in small area estimation, but inverse probability weighting (IPW) gives model-dependent domain estimators whose reliability may vary substantially across domains. Standard Fay-Herriot (FH) smoothing borrows strength across domains, yet it uses the supplied area-level variance estimates as if they fully described the uncertainty of the input estimators. This can be misleading when some domains have weak coverage, unstable weights, or poor auxiliary balance, since these features may indicate selection-bias risk not captured by the estimated variance alone. We propose a diagnostics-guided variance-inflated FH estimator for finite-population domain totals. The method starts from calibrated IPW domain estimators, summarizes their reliability through a small set of domain diagnostics, and introduces a mixture variance-inflation component in the FH observation equation. Domains whose diagnostics indicate weaker IPW information are thereby smoothed more strongly toward the area-level regression mean. A truth-known validation based on a pseudo-real population of Lithuanian business enterprises shows a substantial reduction in estimation error relative to calibrated IPW.

stat.ME↗

Toward design-based inference for data integration

Integrating non-probability samples into finite-population inference typically requires modeling unknown selection probabilities under a missing-at-random (MAR) assumption that is difficult to verify. We propose a design-based alternative in which the non-probability sample is treated as a fully observed certainty stratum and a probability sample is drawn only from the complementary, previously unsampled units. Within this sequential framework, we develop two generalized regression estimators: one fitting the outcome model separately in the complementary stratum, the other pooling both samples; we make two distinct contributions. First, both estimators are design-consistent and admit consistent variance estimators with no assumption whatsoever on the non-probability selection mechanism, including under not-missing-at-random (NMAR) selection. Second, under a working superpopulation model that holds in both strata, the pilot non-probability sample can be used to construct second-stage inclusion probabilities that achieve Isaki-Fuller asymptotic optimality for the separate estimator; this optimality claim relies on assumptions strictly stronger than MAR, but its failure does not invalidate the consistency results above. A diagnostic test for coefficient homogeneity is proposed to guide the choice between the two estimators. Simulations confirm that the sequential estimators remain essentially unbiased under both MAR and NMAR, while propensity-adjusted competitors can be severely biased under NMAR. Two applications from Lithuanian official statistics illustrate that separate regression is preferable when the pilot stratum and its complement are strongly heterogeneous, whereas combined regression offers a modest efficiency gain when the two strata are similar.

stat.ME↗

Small area estimation using incomplete auxiliary information

Auxiliary information is increasingly available from administrative and other data sources, but it is often incomplete and of non-probability origin. We propose a two-step small area estimation approach in which the first step relies on design-based model calibration and exploits a large non-probability source providing a noisy proxy of the study variable for only part of the population. A unit-level measurement-error working model is fitted on the linked overlap between the probability survey and the external source, and its predictions are incorporated through domain-specific model-calibration constraints to obtain approximately design-unbiased domain totals. These totals and their variance estimates are then used in a Fay-Herriot area-level model with exactly known covariates to produce empirical best linear unbiased predictors. The approach is demonstrated in three enterprise survey settings from official statistics by integrating probability sample data with (i) administrative records, (ii) a cut-off data source, and (iii) web-scraped online information. Empirical comparisons show consistent improvements in domain-level precision over direct estimation and over a Fay-Herriot benchmark that directly incorporates the proxy information as an error-prone covariate. These gains are achieved without modeling the selection mechanism of the non-probability sample.

stat.ME↗

Design-based composite estimation of small proportions in small domains

Traditional direct estimation methods are not efficient for domains of a survey population with small sample sizes. To estimate the domain proportions, we combine the direct estimators and the regression-synthetic estimators based on domain-level auxiliary information. For the case of small true proportions, we introduce the design-based linear combination that is a robust alternative to the empirical best linear unbiased predictor (EBLUP) based on the Fay--Herriot model. We also consider an adaptive procedure optimizing a sample-size-dependent composite estimator, which depends on a single parameter for all domains. We imitate the Lithuanian Labor Force Survey, where we estimate the proportions of the unemployed and employed in municipalities. We show where the considered design-based compositions and estimators of their mean square errors are competitive for EBLUP and its accuracy estimation.

stat.ME↗

Design-based composite estimation rediscovered

Small area estimation methods are used in surveys, where sample sizes are too small to get reliable direct estimates of parameters in some population domains. We consider design-based linear combinations of direct and synthetic estimators and propose a two-step procedure to approach the optimal combination. We construct the mean square error estimator suitable for this and any other linear composition that estimates the optimal one. We apply the theory to two design-based compositions analogous to the empirical best linear unbiased predictors (EBLUPs) based on the basic area- and unit-level models. The simulation study shows that the new methods are efficient compared to estimation using EBLUP.

stat.ME↗

A solution in small area estimation problems

We present a new method in problems where estimates are needed for finite population domains with small or even zero sample sizes. In contrast to known estimation methods, an auxiliary information is used to model sizes of population units instead of a direct prediction of their values of interest. In particular, via an additional characterization of regression models, we incorporate a scatter and variabilities of the units sizes into an estimator, and then it uses an information of the whole sample by taking into an account a location of the estimation domain inside the population. To reduce an impact of the introduced domain total estimator bias to the mean square error, we construct also a regression type version of the estimator. An efficiency of the method proposed is shown in a simulation study.

math.ST↗

Gini's mean difference and variance as measures of finite populations scales

We consider Gini's mean difference statistic as an alternative to the empirical variance in the settings of finite populations where simple random samples are drawn without replacement. In particular, we discuss specific (in the finite population context) estimation strategies for a scale of the population, related to the alternative statistic under possible presence of outliers in the data. The paper presents also a wide comparative survey of properties of the Gini mean difference statistic and the empirical variance. It includes asymptotic properties of both statistics: the asymptotic normality, one-term Edgeworth expansions and bootstrap approximations for Studentized versions of the statistics. An estimation of the variances and other parameters of the statistics is also in the study, where we exploit an auxiliary information on the population elements in the case of its availability. Theoretical results are illustrated with a simulation study.

math.ST↗

On the asymptotic normality of finite population L-statistics

We give sufficient conditions for the asymptotic normality of linear combinations of order statistics (L-statistics) in the case of simple random samples without replacement. In the first case, restrictions are imposed on the weights of L-statistics. The second case is on trimmed means, where we introduce a new finite population smoothness condition.

math.ST↗

An Edgeworth expansion for finite population L-statistics

In this paper, we consider the one-term Edgeworth expansion for finite population L-statistics. We provide an explicit formula for the Edgeworth correction term and give sufficient conditions for the validity of the expansion which are expressed in terms of the weight function that defines the statistics and moment conditions.

math.ST↗