SearcharxivSearch

arXiv subjects

Mamadou Yauck

Publications and source records attributed to Mamadou Yauck.

7 recordsLinked to original sources

Estimating hidden population size from a single respondent-driven sampling survey

This work is concerned with the estimation of hard-to-reach population sizes using a single respondent-driven sampling (RDS) survey, a variant of chain-referral sampling that leverages social relationships to reach members of a hidden population. The popularity of RDS as a standard approach for surveying hidden populations brings theoretical and methodological challenges regarding the estimation of population sizes, mainly for public health purposes. This paper proposes a frequentist, model-based framework for estimating the size of a hidden population using a network-based approach. An optimization algorithm is proposed for obtaining the identification region of the target parameter when model assumptions are violated. We characterize the asymptotic behavior of our proposed methodology and assess its finite sample performance under departures from model assumptions.

stat.ME

Small Sample Inference for Two-way Capture Recapture Experiments

The properties of the generalized Waring distribution defined on the non negative integers are reviewed. Formulas for its moments and its mode are given. A construction as a mixture of negative binomial distributions is also presented. Then we turn to the Petersen model for estimating the population size $N$ in a two-way capture recapture experiment. We construct a Bayesian model for $N$ by combining a Waring prior with the hypergeometric distribution for the number of units caught twice in the experiment. Credible intervals for $N$ are obtained using quantiles of the posterior, a generalized Waring distribution. The standard confidence interval for the population size constructed using the asymptotic variance of Petersen estimator and .5 logit transformed interval are shown to be special cases of the generalized Waring credible interval. The true coverage of this interval is shown to be bigger than or equal to its nominal converage in small populations, regardless of the capture probabilities. In addition, its length is substantially smaller than that of the .5 logit transformed interval. Thus the proposed generalized Waring credible interval appears to be the best way to quantify the uncertainty of the Petersen estimator for populations size.

stat.ME

Population Size Estimation for Respondent-Driven Sampling and Capture-Recapture: A Unifying Framework

This paper deals with the estimation of population sizes for respondent-driven sampling (RDS), a variant of link-tracing sampling that leverages social networks over a number of waves to recruit individuals from hidden populations. The RDS process is mostly controlled by individual participants who might report on recruitment proposals, or nominations, that they have received or given. By considering all nominations given or received over a time period, one can create a capture-recapture dataset in which units are individuals who have received at least one nomination and capture occasions are either time intervals or recruitment waves, with the goal of estimating the size $N$ of the hidden population. In this paper, we argue that the underlying process that generated the RDS nomination data is that of a capture-recapture experiment. We then proposed a methodology for the estimation of the population size and investigated its performance against departures from classical capture-recapture assumptions.

stat.ME

On the Estimation of Peer Effects for Sampled Networks

This paper deals with the estimation of exogeneous peer effects for partially observed networks under the new inferential paradigm of design identification, which characterizes the missing data challenge arising with sampled networks with the central idea that two full data versions which are topologically compatible with the observed data may give rise to two different probability distributions. We show that peer effects cannot be identified by design when network links between sampled and unsampled units are not observed. Under realistic modeling conditions, and under the assumption that sampled units report on the size of their network of contacts, the asymptotic bias arising from estimating peer effects with incomplete network data is characterized, and a bias-corrected estimator is proposed. The finite sample performance of our methodology is investigated via simulations.

econ.EM

Neighbourhood Bootstrap for Respondent-Driven Sampling

Respondent-Driven Sampling (RDS) is a form of link-tracing sampling, a sampling technique used for `hard-to-reach' populations that aims to leverage individuals' social relationships to reach potential participants. While the methodological focus has been restricted to the estimation of population proportions, there is a growing interest in the estimation of uncertainty for RDS as recent findings suggest that most variance estimators underestimate variability. Recently, Baraff et al. (2016) proposed the \textit{tree bootstrap} method based on resampling the RDS recruitment tree, and empirically showed that this method outperforms current bootstrap methods. However, some findings suggest that the tree bootstrap (severely) overestimates uncertainty. In this paper, we propose the \textit{neighbourhood} bootstrap method for quantifiying uncertainty in RDS. We prove the consistency of our method under some conditions and investigate its finite sample performance, through a simulation study, under realistic RDS sampling assumptions.

stat.ME

General Regression Methods for Respondent-Driven Sampling Data

Respondent-Driven Sampling (RDS) is a variant of link-tracing sampling techniques that aim to recruit hard-to-reach populations by leveraging individuals' social relationships. As such, an RDS sample has a graphical component which represents a partially observed network of unknown structure. Moreover, it is common to observe homophily, or the tendency to form connections with individuals who share similar traits. Currently, there is a lack of principled guidance on multivariate modeling strategies for RDS to address homophilic covariates and the dependence between observations within the network. In this work, we propose a methodology for general regression techniques using RDS data. This is used to study the socio-demographic predictors of HIV treatment optimism (about the value of antiretroviral therapy) among gay, bisexual and other men who have sex with men, recruited into an RDS study in Montreal, Canada.

stat.ME

Sampling from Networks: Respondent-Driven Sampling

Respondent-Driven Sampling (RDS) is a variant of link-tracing, a sampling technique for surveying hard-to-reach communities that takes advantage of community members' social networks to reach potential participants. As a network-based sampling method, RDS is faced with the fundamental problem of sampling from population networks where features such as homophily (the tendency for individuals with similar traits to share social ties) and differential activity (the ratio of the average number of connections by attribute) are sensitive to the choice of a sampling method. Though not clearly described in the RDS literature, many simple methods exist to generate simulated RDS data, with specific levels of network features, where the focus is on estimating simple estimands. However, the accuracy of these methods in their abilities to consistently recover those targeted network features remains unclear. This is also motivated by recent findings that some population network parameters (e.g.~homophily) cannot be consistently estimated from the RDS data alone \citep{Crawford17}. In this paper, we conduct a simulation study to assess the accuracy of existing RDS simulation methods, in terms of their abilities to generate RDS samples with the desired levels of two network parameters: homophily and differential activity. The results show that (1) homophily cannot be consistently estimated from simulated RDS samples and (2) differential activity estimates are more precise when groups, defined by traits, are equally active and equally represented in the population. We use this approach to mimic features of the Engage Study, an RDS sample of gay, bisexual and other men who have sex with men in Montreal.

stat.AP