arXiv · 2608.26610
On efficiency gains via augmenting a tiny sample with a massive auxiliary sample
Abstract
In this paper, we study the problem of augmenting a tiny target sample with a massive auxiliary sample. Utilizing Tukey's factorization, there are two popular approaches: the inverse probability weight (IPW) and the full-likelihood (FL) methods. We show that the IPW approach suffers from the limited target sample problem while the FL method may estimate some model parameters at the rate of the massive auxiliary sample size, a phenomenon we call full efficiency gain. We study the theory behind the full efficiency gain for exponential families and mixtures of exponential families. We also study the efficiency gain for the IPW method under a nonparametric procedure and show how it can achieve a parametric rate of the target sample size. As a side note, we also discuss how one may use FL to train neural network models simultaneously for both the target distribution and the odds model.
Explore related subjects
Keep this discovery
Yen-Chi Chen. 2026-08-27. On efficiency gains via augmenting a tiny sample with a massive auxiliary sample. https://arxiv.org/abs/2608.26610
Cite the original work for its findings. Save a collection to share your selection of sources.