Searcharxiv⌕ Search

arXiv subjects

Barry I. Graubard

Publications and source records attributed to Barry I. Graubard.

5 recordsLinked to original sources

Nonprobability follow-up sample analysis: an application to SARS-CoV-2 infection prevalence estimation

Public health policy makers are faced with making crucial decisions rapidly during infectious disease outbreaks such as that caused by SARS-CoV-2. Ideally, rapidly deployed representative health surveys could provide needed data for such decisions. Under the constraints of a limited timeframe and resources, it may be infeasible to implement random based (probability) sampling that yields a population representative survey sample with high response rates. As an alternative, a volunteer (nonprobability) sample is often collected using outreach methods such as social media and web surveys. Compared to a probability sample, a nonprobability sample is subject to selection bias. In addition, when participants are followed longitudinally nonresponse often occurs at later follow up timepoints. As a result, estimates of cross-sectional parameters at later timepoints will be subject to selection bias and nonresponse bias. In this paper, we create kernel-weighted pseudoweights (KW) for the baseline survey participants and construct nonresponse-adjusted kw (kwNR) for respondents at each follow-visit to estimate the population mean at the follow-up visits. We develop Taylor Linearization variance estimation that accounts for variability due to estimating both pseudoweights and the nonresponse adjustments. Simulations are conducted to evaluate the proposed kwNR-weighted estimates. We investigate covariate effects on each of the following: baseline sample participation propensity, follow-up response propensity and the mean of the outcome. We apply the proposed kwNR-weighted methods to the SARS-Cov-2 antibody seropositivity longitudinal study, which begins with a baseline survey early in the pandemic, and collects data at six- and twelve-month post baseline follow-ups.

stat.AP↗

Representative Pure Risk Estimation by Using Data from Epidemiologic Studies, Surveys, and Registries: Estimating Risks for Minority Subgroups

Representative risk estimation is fundamental to clinical decision-making. However, risks are often estimated from non-representative epidemiologic studies, which usually underrepresent minorities. "Model-based" methods use population registries to improve externally validity of risk estimation but assume hazard ratios (HR) are generalizable from samples to the target finite population. "Pseudoweighting" methods improve representativeness of studies by using an external probability-based survey as the reference, but the resulting estimators can be biased due to propensity model misspecification or inefficient due to variable pseudoweights or small sample sizes of minorities in the cohort and/or survey. We propose a two-step pseudoweighting procedure that poststratifies the event rates among age/race/sex strata in the pseudoweighted cohort to the population rates to produce efficient and robust pure risk estimation (i.e., a cause-specific absolute risk in the absence of competing events). For developing an all-cause mortality risk model representative for the US, our findings suggest that HRs for minorities are not generalizable, and that surveys can have inadequate numbers of events for minorities. Poststratification on event rates is crucial for obtaining reliable risk estimation for minority subgroups.

stat.ME↗

Efficient and Robust Propensity-Score-Based Methods for Population Inference using Epidemiologic Cohorts

Most epidemiologic cohorts are composed of volunteers who do not represent the general population. To enable population inference from cohorts, we and others have proposed utilizing probability survey samples as external references to develop a propensity score (PS) for membership in the cohort versus survey. Herein we develop a unified framework for PS-based weighting (such as inverse PS weighting (IPSW)) and matching methods (such as kernel-weighting (KW) method). We identify a fundamental Strong Exchangeability Assumption (SEA) underlying existing PS-based matching methods whose failure invalidates inference even if the PS-model is correctly specified. We relax the SEA to a Weak Exchangeability Assumption (WEA) for the matching method. Also, we propose IPSW.S and KW.S methods that reduce the variance of PS-based estimators by scaling the survey weights used in the PS estimation. We prove consistency of the IPSW.S and KW.S estimators of population means and prevalences under WEA, and provide asymptotic variances and consistent variance estimators. In simulations, the KW.S and IPSW.S estimators had smallest MSE. In our data example, the original KW estimates had large bias, whereas the KW.S estimates had the smallest MSE.

stat.ME↗

Semiparametric regression in testicular germ cell data

It is possible to approach regression analysis with random covariates from a semiparametric perspective where information is combined from multiple multivariate sources. The approach assumes a semiparametric density ratio model where multivariate distributions are "regressed" on a reference distribution. A kernel density estimator can be constructed from many data sources in conjunction with the semiparametric model. The estimator is shown to be more efficient than the traditional single-sample kernel density estimator, and its optimal bandwidth is discussed in some detail. Each multivariate distribution and the corresponding conditional expectation (regression) of interest are estimated from the combined data using all sources. Graphical and quantitative diagnostic tools are suggested to assess model validity. The method is applied in quantifying the effect of height and age on weight of germ cell testicular cancer patients. Comparisons are made with multiple regression, generalized additive models (GAM) and nonparametric kernel regression.

stat.ME↗

A Conversation with Monroe Sirken

Born January 11, 1921 in New York City, Monroe Sirken grew up in a suburb of Pasadena, California. He earned B.A. and M.A. degrees in sociology at UCLA in 1946 and 1947, and a Ph.D. in 1950 in sociology with a minor in mathematics at the University of Washington in 1950 where Professor Z. W. Birnbaum was his mentor and thesis advisor. As a Post-Doctoral Fellow of the Social Science Research Council, Monroe spent 1950--1951 at the Statistics Laboratory, University of California at Berkeley and the Office of the Assistant Director for Research, U.S. Bureau of the Census in Suitland, Maryland. Monroe visited the Census Bureau at a time of great change in the use of sampling and survey methods, and decided to remain. He began his government career there in 1951 as a mathematical statistician, and moved to the National Office of Vital Statistics (NOVS) in 1953 where he was an actuarial mathematician and a mathematical statistician. He has held a variety of research and administrative positions at the National Center for Health Statistics (NCHS) and he was the Associate Director, Research and Methodology and the Director, Office of Research and Methodology until 1996 when he became a senior research scientist, the title he currently holds. Aside from administrative responsibilities, Monroe's major professional interests have been conducting and fostering survey and statistical research responsive to the needs of federal statistics. His interest in the design of rare and sensitive population surveys led to the development of network sampling which improves precision by linking multiple selection units to the same observation units. His interest in fostering research on the cognitive aspects of survey methods led to the establishment of permanent questionnaire design research laboratories, first at NCHS and later at other federal statistical agencies here and abroad.

stat.ME↗