SearcharxivSearch

arXiv subjects

Dennis M. Feehan

Publications and source records attributed to Dennis M. Feehan.

7 recordsLinked to original sources

Births are difficult to predict even with rich survey and full-population register data

Major life events have proven difficult to predict. Does this reflect limits of theory, data, and algorithms, or the large role of chance? We examine one outcome - having a child within three years - through a near-ideal setting for prediction: a data challenge where 147 researchers predicted births for Dutch residents aged 18-45, using survey data and full-population registers. Methods ranged from logistic regression to a large language model and transformers. Predictions were moderately accurate (best F1: register 0.59, survey 0.76); advanced models did not outperform classical ones; and the larger registers did not beat the survey. Simulating the stochastic biology of conception and pregnancy, we estimated a predictive ceiling (survey F1 ~ 0.86-0.94, register 0.88-0.96). Observed performance falls short of this ceiling, implicating imperfect data, methods, and unmodelled chance, while the ceiling itself shows that chance in reproduction alone sets a non-trivial limit on predicting individual lives.

cs.LG

A sample size heuristic for network scale-up studies

The network scale-up method (NSUM) is a survey-based method for estimating the number of individuals in a hidden or hard-to-reach subgroup of a general population. In NSUM surveys, sampled individuals report how many others they know in the subpopulation of interest (e.g. "How many sex workers do you know?") and how many others they know in subpopulations of the general population (e.g. "How many bus drivers do you know?"). NSUM is widely used to estimate the size of important epidemiological risk groups, including men who have sex with men, sex workers, HIV+ individuals, and drug users. Unlike several other methods for population size estimation, NSUM requires only a single random sample and the estimator has a conveniently simple form. Despite its popularity, there are no published guidelines for the minimum sample size calculation to achieve a desired statistical precision. Here, we provide a sample size formula that can be employed in any NSUM survey. We show analytically and by simulation that the sample size controls error at the nominal rate and is robust to some forms of network model mis-specification. We apply this methodology to study the minimum sample size and relative error properties of several published NSUM surveys.

stat.ME

Using an online sample to learn about an offline population

Online data sources offer tremendous promise to demography and other social sciences, but researchers worry that the group of people who are represented in online datasets can be different from the general population. We show that by sampling and anonymously interviewing people who are online, researchers can learn about both people who are online and people who are offline. Our approach is based on the insight that people everywhere are connected through in-person social networks, such as kin, friendship, and contact networks. We illustrate how this insight can be used to derive an estimator for tracking the *digital divide* in access to the internet, an increasingly important dimension of population inequality in the modern world. We conducted a large-scale empirical test of our approach, using an online sample to estimate internet adoption in five countries ($n \approx 15,000$). Our test embedded a randomized experiment whose results can help design future studies. Our approach could be adapted to many other settings, offering one way to overcome some of the major challenges facing demographers in the information age.

stat.AP

Estimating adult death rates from sibling histories: A network approach

Hundreds of millions of people live in countries that do not have complete death registration systems, meaning that most deaths are not recorded and critical quantities like life expectancy cannot be directly measured. The sibling survival method is a leading approach to estimating adult mortality in the absence of death registration. The idea is to ask a survey respondent to enumerate her siblings and to report about their survival status. In many countries and time periods, sibling survival data are the only nationally-representative source of information about adult mortality. Although a huge amount of sibling survival data has been collected, important methodological questions about the method remain unresolved. To help make progress on this issue, we propose re-framing the sibling survival method as a network sampling problem. This approach enables us to formally derive statistical estimators for sibling survival data. Our derivation clarifies the precise conditions that sibling history estimates rely upon; it leads to internal consistency checks that can help assess data and reporting quality; and it reveals important quantities that could potentially be measured to relax assumptions in the future. We introduce the R package siblingsurvival, which implements the methods we describe.

stat.AP

Separating the signal from the noise: Evidence for deceleration in old-age death rates

Widespread population aging has made it critical to understand death rates at old ages. However, studying mortality at old ages is challenging because the data are sparse: numbers of survivors and deaths get smaller and smaller with age. We show how to address this challenge by using principled model selection techniques to empirically evaluate theoretical mortality models. We test nine different theoretical models of old-age death rates by fitting them to 360 high-quality datasets on cohort mortality above age 80. Models that allow for the possibility of decelerating death rates tend to fit better than models that assume exponentially increasing death rates. No single model is capable of universally explaining observed old-age mortality patterns, but the Log-Quadratic model most consistently predicts well. Patterns of model fit differ by country and sex; we discuss possible mechanisms, including sample size, period effects, and regional or cultural factors that may be important keys to understanding patterns of old-age mortality. We introduce a freely available R package that enables researchers to extend our analysis to other models, age ranges, and data sources.

stat.AP

The network survival method for estimating adult mortality: Evidence from a survey experiment in Rwanda

Adult death rates are a critical indicator of population health and wellbeing. Wealthy countries have high-quality vital registration systems, but poor countries lack this infrastructure and must rely on estimates that are often problematic. In this paper, we introduce the network survival method, a new approach for estimating adult death rates. We derive the precise conditions under which it produces estimates that are consistent and unbiased. Further, we develop an analytical framework for sensitivity analysis. To assess the performance of the network survival method in a realistic setting, we conducted a nationally-representative survey experiment in Rwanda (n=4,669). Network survival estimates were similar to estimates from other methods, even though the network survival estimates were made with substantially smaller samples and are based entirely on data from Rwanda, with no need for model life tables or pooling of data from other countries. Our analytic results demonstrate that the network survival method has attractive properties, and our empirical results show that it can be used in countries where reliable estimates of adult death rates are sorely needed.

stat.AP

Generalizing the Network Scale-Up Method: A New Estimator for the Size of Hidden Populations

The network scale-up method enables researchers to estimate the size of hidden populations, such as drug injectors and sex workers, using sampled social network data. The basic scale-up estimator offers advantages over other size estimation techniques, but it depends on problematic modeling assumptions. We propose a new generalized scale-up estimator that can be used in settings with non-random social mixing and imperfect awareness about membership in the hidden population. Further, the new estimator can be used when data are collected via complex sample designs and from incomplete sampling frames. However, the generalized scale-up estimator also requires data from two samples: one from the frame population and one from the hidden population. In some situations these data from the hidden population can be collected by adding a small number of questions to already planned studies. For other situations, we develop interpretable adjustment factors that can be applied to the basic scale-up estimator. We conclude with practical recommendations for the design and analysis of future studies.

stat.AP