SearcharxivSearch

arXiv subjects

Niamh Cahill

Publications and source records attributed to Niamh Cahill.

16 recordsLinked to original sources

A Scalable Bayesian Spatiotemporal Model for Water Level Predictions using a Nearest Neighbor Gaussian Process Approach

Obtaining accurate water level predictions are essential for water resource management and implementing flood mitigation strategies. Several data-driven models can be found in the literature. However, there has been limited research with regard to addressing the challenges posed by large spatio-temporally referenced hydrological datasets, in particular, the challenges of maintaining predictive performance and uncertainty quantification. Gaussian Processes (GPs) are commonly used to capture complex space-time interactions. However, GPs are computationally expensive and suffer from poor scaling as the number of locations increases due to required covariance matrix inversions. To overcome the computational bottleneck, the Nearest Neighbor Gaussian Process (NNGP) introduces a sparse precision matrix providing scalability without having to make inferential compromises. In this work we introduce an innovative model in the hydrology field, specifically designed to handle large datasets consisting of a large number of spatial points across multiple hydrological basins, with daily observations over an extended period. We investigate the application of a Bayesian spatiotemporal NNGP model to a rich dataset of daily water levels of rivers located in Ireland. The dataset comprises a network of 301 stations situated in various basins across Ireland, measured over a period of 90 days. The proposed approach allows for prediction of water levels at future time points, as well as the prediction of water levels at unobserved locations through spatial interpolation, while maintaining the benefits of the Bayesian approach, such as uncertainty propagation and quantification. Our findings demonstrate that the proposed model outperforms competing approaches in terms of accuracy and precision.

stat.AP

Using Time Series Measures to Explore Family Planning Survey Data and Model-based Estimates

Family planning is a global development priority and a key indicator of reproductive health. Monitoring progress is challenged by gaps in survey data across countries. The United Nations Population Division addresses this with the Family Planning Estimation Model (FPEM), a Bayesian hierarchical time series model producing annual estimates of modern contraceptive use while sharing information across countries and regions. This paper evaluates how well FPEM estimates align with survey data using time series diagnostic indices from the wdiexplorer R package, which account for countries nested within sub-regions. Visualisation of survey data, modelled trajectories, and diagnostics enables assessment of model performance, highlighting where trends align and where discrepancies occur.

stat.AP

Bayesian Statistical Modeling in Action for Estimation and Forecasting in Low- and Middle-income Countries: The Case of the Family Planning Estimation Tool

The Family Planning Estimation Tool (FPET) is used in low- and middle-income countries to produce estimates and short-term forecasts of family planning indicators, such as modern contraceptive use and unmet need for contraceptives. Estimates are obtained via a Bayesian statistical model that is fitted to country-specific data from surveys and service statistics data. The model has evolved over the last decade based on user inputs. In this paper we summarize the main features of the statistical model used in FPET and introduce recent updates related to capturing contraceptive transitions, fitting to survey data that may be error prone, and the use of service statistics data. We assess model performance through a validation exercise and find that FPET is reasonably well calibrated. We use our experience with FPET to briefly discuss lessons learned and open challenges related to the broader field of statistical modeling for monitoring of demographic and global health indicators.

stat.AP

wdiexplorer: An R package Designed for Exploratory Analysis of World Development Indicators (WDI) Data

The World Development Indicators (WDI) database provides a wide range of global development data, maintained and published by the World Bank. Our \textit{wdiexplorer} package offers a comprehensive workflow that sources WDI data via the \textit{WDI} R package, prepares and explores country-level panel data of the WDI through computational functions to calculate diagnostic metrics and visualise the outputs. By leveraging the functionalities of \textit{wdiexplorer} package, users can efficiently explore any indicator dataset of the WDI, compute diagnostic indices, and visualise the metrics by incorporating the pre-defined grouping structures to identify patterns, outliers, and other interesting features of temporal behaviours. This paper presents the \textit{wdiexplorer} package, demonstrates its functionalities using the WDI: PM$_{2.5}$ air pollution dataset, and discusses the observed patterns and outliers across countries and within groups of country-level panel data.

stat.CO

Bayesian probabilistic projections of proportions with limited data: An application to subnational contraceptive method supply shares

Engaging the private sector in contraceptive method supply is critical for creating equitable, sustainable, and accessible healthcare systems. To achieve this, it is essential to understand where women obtain their modern contraceptives. While national-level estimates provide valuable insights into overall trends in contraceptive supply, they often obscure variation within and across subnational regions. Addressing localized needs has become increasingly important as countries adopt decentralized models for family planning services. Decentralization has also underscored the need for reliable subnational estimates of key family planning indicators. The absence of regularly collected subnational data has hindered effective monitoring and decision-making. To bridge this gap, we propose a novel approach that leverages latent attributes in Demographic and Health Survey (DHS) data to produce Bayesian probabilistic projections of contraceptive method supply shares (the proportions of modern contraceptive methods supplied by public and private sectors) with limited data. Our modeling framework is built on Bayesian hierarchical models. Using penalized splines to track public and private supply shares over time, we leverage the spatial nature of the data and incorporate a correlation structure between recent supply share observations at national and subnational levels. This framework contributes to the domain of subnational estimation of proportions in data-sparse settings, outperforming comparable and previous approaches. As decentralization continues to reshape family planning services, producing reliable subnational estimates of key indicators is increasingly vital for researchers and policymakers.

stat.ME

Enhancing the use of family planning service statistics using a Bayesian modelling approach to inform estimates of modern contraceptive use in low- and middle-income countries

Monitoring family planning indicators, such as modern contraceptive prevalence rate (mCPR), is essential for family planning programming. The Family Planning Estimation Tool (FPET) uses survey data to estimate and forecast family planning indicators, including mCPR, over time. However, sole reliance on large-scale surveys, carried out on average every 3-5 years, can lead to data gaps. Service statistics are a readily available data source, routinely collected in conjunction with service delivery. Various service statistics data types can be used to derive a family planning indicator called Estimated Modern Use (EMU). In a number of countries, annual rates of change in EMU have been found to be predictive of true rates of change in mCPR. However, it has been challenging to capture the varying levels of uncertainty associated with the EMU indicator across different countries and service statistics data types and to subsequently quantify this uncertainty when using EMU in FPET. We present a new approach to using EMUs in FPET to inform mCPR estimates, using annual EMU rates of change as input, and accounting for uncertainty associated with the EMU derivation process. The approach also considers additional country-type-specific uncertainty. We assess the EMU type-specific uncertainty at the country level, via a Bayesian hierarchical modelling approach. Validation results and anonymised country-level case studies highlight improved predictive performance and provide insights into the impact of including EMU data on mCPR estimates compared to using survey data alone. Together, they demonstrate that EMUs can help countries monitor progress toward their family planning goals more effectively.

stat.AP

Visualisation for Exploratory Modelling Analysis of Bayesian Hierarchical Models

When developing Bayesian hierarchical models, selecting the most appropriate hierarchical structure can be a challenging task, and visualisation remains an underutilised tool in this context. In this paper, we consider visualisations for the display of hierarchical models in data space and compare a collection of multiple models via their parameters and hyper-parameter estimates. Specifically, with the aim of aiding model choice, we propose new visualisations to explore how the choice of Bayesian hierarchical modelling structure impacts parameter distributions. The visualisations are designed using a robust set of principles to provide richer comparisons that extend beyond the conventional plots and numerical summaries typically used. As a case study, we investigate five Bayesian hierarchical models fit using the brms R package, a high-level interface to Stan for Bayesian modelling, to model country mathematics trends from the PISA (Programme for International Student Assessment) database. Our case study demonstrates that by adhering to these principles, researchers can create visualisations that not only help them make more informed choices between Bayesian hierarchical model structures but also enable them to effectively communicate the rationale for those choices.

stat.ME

mcmsupply: An R Package for Estimating Contraceptive Method Market Supply Shares

In this paper, we introduce the R package mcmsupply which implements Bayesian hierarchical models for estimating and projecting modern contraceptive market supply shares over time. The package implements four model types. These models vary by the administration level of their outcome estimates (national or subnational estimates) and dataset type utilised in the estimation (multi-country or single-country contraceptive market supply datasets). mcmsupply contains a compilation of national and subnational level contraceptive source datasets, generated by IPUMS and Demographic and Health Survey microdata. We describe the functions that implement the models through practical examples. The annual estimates and projections with uncertainty of the contraceptive market supply, produced by mcmcsupply at a national and subnational level, are the first of their kind. These estimates and projections have diverse applications, including acting as an indicator of family planning market stability over time and being utilised in the calculation of estimates of modern contraceptive use.

stat.ME

Diagnostics for categorical response models based on quantile residuals and distance measures

Polytomous categorical data are frequent in studies, that can be obtained with an individual or grouped structure. In both structures, the generalized logit model is commonly used to relate the covariates on the response variable. After fitting a model, one of the challenges is the definition of an appropriate residual and choosing diagnostic techniques. Since the polytomous variable is multivariate, raw, Pearson, or deviance residuals are vectors and their asymptotic distribution is generally unknown, which leads to difficulties in graphical visualization and interpretation. Therefore, the definition of appropriate residuals and the choice of the correct analysis in diagnostic tools is important, especially for nominal data, where a restriction of methods is observed. This paper proposes the use of randomized quantile residuals associated with individual and grouped nominal data, as well as Euclidean and Mahalanobis distance measures, as an alternative to reduce the dimension of the residuals. We developed simulation studies with both data structures associated. The half-normal plots with simulation envelopes were used to assess model performance. These studies demonstrated a good performance of the quantile residuals, and the distance measurements allowed a better interpretation of the graphical techniques. We illustrate the proposed procedures with two applications to real data.

stat.ME

reslr: An R package for relative sea level modelling

We present reslr, an R package to perform Bayesian modelling of relative sea level data. We include a variety of different statistical models previously proposed in the literature, with a unifying framework for loading data, fitting models, and summarising the results. Relative sea-level data often contain measurement error in multiple dimensions and so our package allows for these to be included in the statistical models. When plotting the output sea level curves, the focus is often on comparing rates of change, and so our package allows for computation of the derivative of sea level curves with appropriate consideration of the uncertainty. We provide a large example dataset from the Atlantic coast of North America and show some of the results that might be obtained from our package.

stat.AP

A noisy-input generalised additive model for relative sea-level change along the Atlantic coast of North America

We propose a Bayesian, noisy-input, spatial-temporal generalised additive model to examine regional relative sea-level (RSL) changes over time. The model provides probabilistic estimates of component drivers of regional RSL change via the combination of a univariate spline capturing a common regional signal over time, random slopes and intercepts capturing site-specific (local), long-term linear trends and a spatial-temporal spline capturing residual, non-linear, local variations. Proxy and instrumental records of RSL and corresponding measurement errors inform the model and a noisy-input method accounts for proxy temporal uncertainties. Results focus on the decomposition of RSL over the past 3000 years along the Atlantic coast of North America.

stat.AP

Estimating the proportion of modern contraceptives supplied by the public and private sectors using a Bayesian hierarchical penalized spline model

Quantifying the public/private sector supply of contraceptive methods within countries is vital for effective and sustainable family planning (FP) delivery. In many low and middle-income countries (LMIC), measuring the contraceptive supply source often relies on Demographic Health Surveys (DHS). However, many of these countries carry out the DHS approximately every 3-5 years and do not have recent data beyond 2015/16. Our objective in estimating the set of related contraceptive supply-share outcomes (proportion of modern contraceptive methods supplied by the public/private sectors) is to take advantage of latent attributes present in dataset to produce annual, country-specific estimates and projections with uncertainty. We propose a Bayesian, hierarchical, penalized-spline model with multivariate-normal spline coefficients to capture cross-method correlations. Our approach offers an intuitive way to share information across countries and sub-continents, model the changes in the contraceptive supply share over time, account for survey observational errors and produce probabilistic estimates and projections that are informed by past changes in the contraceptive supply share as well as correlations between rates of change across different methods. These results will provide valuable information for evaluating FP program effectiveness. To the best of our knowledge, it is the first model of its kind to estimate these quantities.

stat.ME

A Bayesian Hierarchical Time Series Model for Reconstructing Hydroclimate from Multiple Proxies

We propose a Bayesian hierarchical model which produces probabilistic reconstructions of hydroclimatic variability in Queensland Australia. The model provides a standardised approach to hydroclimate reconstruction using multiple palaeoclimate proxy records derived from natural archives such as speleothems, ice cores and tree rings. The method combines time-series modelling with inverse prediction to quantify the relationships between a given hydroclimate index and relevant proxies over an instrumental period and subsequently reconstruct the hydroclimate back through time. We present case studies for Brisbane and Fitzroy catchments focusing on two hydroclimate indices, the Rainfall Index (RFI) and the Standardised Precipitation-Evapotranspiration Index (SPEI). The probabilistic nature of the reconstructions allows us to estimate the probability that a hydroclimate index in any reconstruction year was lower (higher) than the minimum (maximum) value observed over the instrumental period. In Brisbane, the RFI is unlikely (probabilities < 20%) to have exhibited extremes beyond the minimum/maximum values observed between 1889 and 2017. However, in Fitzroy there are several years during the reconstruction period where the RFI is likely (> 50% probability) to have exhibited behaviour beyond the minimum/maximum of what has been observed. For SPEI, the probability of observing such extremes since the end of the instrumental period in 1889 doesn't exceed 50% in any reconstruction year in Brisbane or Fitzroy.

stat.AP

Statistical modeling of rates and trends in Holocene relative sea level

Characterizing the spatio-temporal variability of relative sea level (RSL) and estimating local, regional, and global RSL trends requires statistical analysis of RSL data. Formal statistical treatments, needed to account for the spatially and temporally sparse distribution of data and for geochronological and elevational uncertainties, have advanced considerably over the last decade. Time-series models have adopted more flexible and physically-informed specifications with more rigorous quantification of uncertainties. Spatio-temporal models have evolved from simple regional averaging to frameworks that more richly represent the correlation structure of RSL across space and time. More complex statistical approaches enable rigorous quantification of spatial and temporal variability, the combination of geographically disparate data, and the separation of the RSL field into various components associated with different driving processes. We review the range of statistical modeling and analysis choices used in the literature, reformulating them for ease of comparison in a common hierarchical statistical framework. The hierarchical framework separates each model into different levels, clearly partitioning measurement and inferential uncertainty from process variability. Placing models in a hierarchical framework enables us to highlight both the similarities and differences among modeling and analysis choices. We illustrate the implications of some modeling and analysis choices currently used in the literature by comparing the results of their application to common datasets within a hierarchical framework. In light of the complex patterns of spatial and temporal variability exhibited by RSL, we recommend non-parametric approaches for modeling temporal and spatio-temporal RSL.

stat.AP

Modeling sea-level change using errors-in-variables integrated Gaussian processes

We perform Bayesian inference on historical and late Holocene (last 2000 years) rates of sea-level change. The input data to our model are tide-gauge measurements and proxy reconstructions from cores of coastal sediment. These data are complicated by multiple sources of uncertainty, some of which arise as part of the data collection exercise. Notably, the proxy reconstructions include temporal uncertainty from dating of the sediment core using techniques such as radiocarbon. The model we propose places a Gaussian process prior on the rate of sea-level change, which is then integrated and set in an errors-in-variables framework to take account of age uncertainty. The resulting model captures the continuous and dynamic evolution of sea-level change with full consideration of all sources of uncertainty. We demonstrate the performance of our model using two real (and previously published) example data sets. The global tide-gauge data set indicates that sea-level rise increased from a rate with a posterior mean of 1.13 mm$/$yr in 1880 AD (0.89 to 1.28 mm$/$yr 95% credible interval for the posterior mean) to a posterior mean rate of 1.92 mm$/$yr in 2009 AD (1.84 to 2.03 mm$/$yr 95% credible interval for the posterior mean). The proxy reconstruction from North Carolina (USA) after correction for land-level change shows the 2000 AD rate of rise to have a posterior mean of 2.44 mm$/$yr (1.91 to 3.01 mm$/$yr 95% credible interval). This is unprecedented in at least the last 2000 years.

stat.AP

A Bayesian Hierarchical Model for Reconstructing Sea Levels: From Raw Data to Rates of Change

We present a holistic Bayesian hierarchical model for reconstructing the continuous and dynamic evolution of relative sea-level (RSL) change with fully quantified uncertainty. The reconstruction is produced from biological (foraminifera) and geochemical (δ13C) sea-level indicators preserved in dated cores of salt-marsh sediment. Our model is comprised of three modules: (1) A Bayesian transfer function for the calibration of foraminifera into tidal elevation, which is flexible enough to formally accommodate additional proxies (in this case bulk-sediment δ13C values); (2) A chronology developed from an existing Bchron age-depth model, and (3) An existing errors-in-variables integrated Gaussian process (EIV-IGP) model for estimating rates of sea-level change. We illustrate our approach using a case study of Common Era sea-level variability from New Jersey, U.S.A. We develop a new Bayesian transfer function (B-TF), with and without the δ13C proxy and compare our results to those from a widely-used weighted-averaging transfer function (WA-TF). The formal incorporation of a second proxy into the B-TF model results in smaller vertical uncertainties and improved accuracy for reconstructed RSL. The vertical uncertainty from the multi-proxy B-TF is ~28% smaller on average compared to the WA-TF. When evaluated against historic tide-gauge measurements, the multi-proxy B-TF most accurately reconstructs the RSL changes observed in the instrumental record (MSE = 0.003). The holistic model provides a single, unifying framework for reconstructing and analysing sea level through time. This approach is suitable for reconstructing other paleoenvironmental variables using biological proxies.

stat.AP