SearcharxivSearch

arXiv subjects

Runjia Zou

Publications and source records attributed to Runjia Zou.

3 recordsLinked to original sources

Marginal generalized raking with parametric working models

Generalized raking (GR) was originally developed in the survey statistics literature to incorporate auxiliary information in estimation. Recently, it has been used in the biostatistical and epidemiological literature to estimate regression coefficients in parametric models in cases with missing data, including missing data by design (e.g., two-phase studies). In the regression parameter context, the optimal GR estimator has been shown to be equivalent to the optimal augmented inverse probability weighted estimator. In this paper, we generalize the influence function-based theory for GR to marginal estimands; we call our approach \textit{marginal generalized raking}. We compare our approach to a naive procedure that marginalizes a conditional GR estimator of regression parameters in both fully-synthetic simulations and in an application using data from an observational cohort of persons living with HIV.

stat.ME

Testing hypotheses via orthogonalization

Classical hypothesis testing frameworks break down in contemporary settings in which null hypotheses are increasingly abstract, the same data are used to both generate and test hypotheses, and minimal assumptions about the underlying data are made. In this work, we propose a new framework for conducting valid hypothesis tests in broad contexts. We propose to add and subtract external noise generated from a symmetric shift-family to our data, $X$, to partition it into two pieces, $X^{(1)}$ and $X^{(2)}$. We provide a generic strategy for orthogonalizing $X^{(2)}$ against $X^{(1)}$ under the null hypothesis $H_0$, then show that testing whether the orthogonalization was successful provides a valid test of $H_0$ under mild assumptions. Remarkably, this framework extends naturally to the post-selection inference setting: we simply select a hypothesis on $X^{(1)}$, then perform orthogonalization under the selected null. As our approach neither requires pre-specification of the selection mechanism, nor is restricted to a small class of data-generating distributions, it dramatically expands the settings for which valid post-selection inference can be conducted. We showcase the flexibility of our proposal in several case studies involving challenging pre-specified null hypotheses and post-selection inference scenarios.

stat.ME

Generalized Prediction-Powered Inference, with Application to Binary Classifier Evaluation

In the partially-observed outcome setting, a recent set of proposals known as "prediction-powered inference" (PPI) involve (i) applying a pre-trained machine learning model to predict the response, and then (ii) using these predictions to obtain an estimator of the parameter of interest with asymptotic variance no greater than that which would be obtained using only the labeled observations. While existing PPI proposals consider estimators arising from M-estimation, in this paper we generalize PPI to any regular asymptotically linear estimator. Furthermore, by situating PPI within the context of an existing rich literature on missing data and semi-parametric efficiency theory, we show that while PPI does not achieve the semi-parametric efficiency lower bound outside of very restrictive and unrealistic scenarios, it can be viewed as a computationally-simple alternative to proposals in that literature. We exploit connections to that literature to propose modified PPI estimators that can handle three distinct forms of covariate distribution shift. Finally, we illustrate these developments by constructing PPI estimators of true positive rate, false positive rate, and area under the curve via numerical studies.

stat.ME