SearcharxivSearch

arXiv subjects

Jacob Parsons

Publications and source records attributed to Jacob Parsons.

3 recordsLinked to original sources

A Bayesian hierarchical modeling approach to combining multiple data sources: A case study in size estimation

To combat the HIV/AIDS pandemic effectively, targeted interventions among certain key populations play a critical role. Examples of such key populations include sex workers, people who inject drugs, and men who have sex with men. While having accurate estimates for the size of these key populations is important, any attempt to directly contact or count members of these populations is difficult. As a result, indirect methods are used to produce size estimates. Multiple approaches for estimating the size of such populations have been suggested but often give conflicting results. It is therefore necessary to have a principled way to combine and reconcile these estimates. To this end, we present a Bayesian hierarchical model for estimating the size of key populations that combines multiple estimates from different sources of information. The proposed model makes use of multiple years of data and explicitly models the systematic error in the data sources used. We use the model to estimate the size of people who inject drugs in Ukraine. We evaluate the appropriateness of the model and compare the contribution of each data source to the final estimates.

stat.AP

Evaluating the relative contribution of data sources in a Bayesian analysis with the application of estimating the size of hard to reach populations

When using multiple data sources in an analysis, it is important to understand the influence of each data source on the analysis and the consistency of the data sources with each other and the model. We suggest the use of a retrospective value of information framework in order to address such concerns. Value of information methods can be computationally difficult. We illustrate the use of computational methods that allow these methods to be applied even in relatively complicated settings. In illustrating the proposed methods, we focus on an application in estimating the size of hard to reach populations. Specifically, we consider estimating the number of injection drug users in Ukraine by combining all available data sources spanning over half a decade and numerous sub-national areas in the Ukraine. This application is of interest to public health researchers as this hard to reach population that plays a large role in the spread of HIV. We apply a Bayesian hierarchical model and evaluate the contribution of each data source in terms of absolute influence, expected influence, and level of surprise. Finally we apply value of information methods to inform suggestions on future data collection.

stat.AP

The Value of Information in Retrospect

In the course of any statistical analysis, it is necessary to consider issues of data quality and model appropriateness. Value of information methods were initially put forward in the middle of the twentieth century in order to provide a framework for choosing between potential sources of information. However, since their genesis, value of information methods have been largely neglected by statisticians. In this paper we review and extend existing value of information methods and recommend the use of three quantities for identifying influential and outlying data: an influence measure previously suggested by \cite{kempthorne1986}, a related quantity known as the expected value of sample information that is used to gauge how much influence we would expect a portion of the data to have, and the ratio of these two quantities which serves as a comparison between observed influence and expected influence. We study the basic theoretical properties of those quantities and illustrate our proposed approach using two datasets. A data set containing employment rates and other economic factors in U.S. first presented by \cite{longley} is used to provide an example in the case of linear regression. HIV surveillance data collected from prenatal clinics have been the main source of information for monitoring the HIV epidemic in low and middle income countries. A data set providing information about HIV prevalence in Swaziland is used as an example in the case of generalized linear mixed models.

stat.ME