SearcharxivSearch

arXiv subjects

Samuel I. Watson

Publications and source records attributed to Samuel I. Watson.

6 recordsLinked to original sources

Embracing Spillover: Spatial Effects in Experiments

Interventions delivered in space generate effects that spill over between experimental units. We develop a framework for spatial experiments in which causal estimands, inclduing direct, indirect, total, and dose--response effects, are linear functionals of a spillover kernel. Under cluster randomisation with a fixed exposure set size, the data identify the shape of the kernel but not its level: the level is aliased with the intercept and direct effect, at any sample size and under any outcome model. Every estimator therefore rests on an anchoring assumption that fixes the level. We derive an exact decomposition of the bias of any anchored estimator into a level term, a shape error from kernel misspecification, and a leakage term absorbed by the realised geometry, each computable from the design before data collection. A single scalar, the level leverage, gives each estimand's exposure to the unidentified level, the exact level bias, and the variance cost of estimating the level instead. The conventional cluster-trial analysis is the special case of an implicit anchor, with contamination bias in closed form. Existing approaches, including identification through Bernoulli randomisation, elicited bounds on interference decay, and assumed compact support, are, within the linear exposure mapping, anchoring choices in this framework. We give a taxonomy of design augmentations that purchase the level and a criterion for when to augment and when to anchor.

stat.ME

Two-stage Adaptive Design Cluster Randomised Trials

Adaptive sample size re-estimation, early stopping, and trial re-design at interim analyses can reduce expected sample sizes in randomised trials. Cluster randomised trials, in which groups of participants are randomly allocated to treatment status, may particularly benefit as they can be costly and their required sample sizes depend on one or more auxiliary parameters governing correlations within and between clusters, which are often estimated with high uncertainty. We adapt a combination test approach to the cluster trial setting allowing for early stopping for futility or efficacy and accounting for correlations between trial stages and other nuisance parameters. We consider design decisions for multi-dimensional sample sizes involving clusters, participants, and time and allowing for modifications to intervention roll-out patterns. We use a Pareto optimality approach to balance objectives relating to different components of the sample size and costs. We also examine the interim estimation of auxiliary parameters and trial re-design for efficiency. We illustrate the methods including an example of a parallel cluster trial re-design and a re-analysis of the large cluster randomised trial E-MOTIVE.

stat.ME

Approximate Likelihood-Based Inference for Spatial Generalized Linear Mixed Models

We study maximum likelihood estimation for spatial generalized linear mixed models with Gaussian process approximations using a stochastic Newton-Raphson algorithm. We consider two Gaussian Process approximations in this context: spectral Gaussian process approximations and stochastic partial differential equations (SPDE). We refine the stochastic maximum likelihood algorithm and we propose a new stopping criterion for efficient termination to prevent long runs of sampling in the stationary post-convergence phase and a Monte Carlo estimator of fixed effect standard errors. We run a series of simulation comparisons of spatial statistical models alongside the popular Bayesian integrated nested Laplacian approximation method which incorporates SPDE. We show that HSGP provides nominal coverage of fixed and random effect parameters with smooth latent fields but performance degrades for rough fields. SPDE in a stochastic maximum likelihood framework maintains nominal coverage and matches or improves upon the performance of Bayesian integrated nested Laplacian approximation.

stat.ME

Generalised Linear Mixed Model Specification, Analysis, Fitting, and Optimal Design in R with the glmmr Packages

We describe the \proglang{R} package \pkg{glmmrBase} and an extension \pkg{glmmrOptim}. \pkg{glmmrBase} provides a flexible approach to specifying, fitting, and analysing generalised linear mixed models. We use an object-orientated class system within \proglang{R} to provide methods for a wide range of covariance and mean functions, including specification of non-linear functions of data and parameters, relevant to multiple applications including cluster randomised trials, cohort studies, spatial and spatio-temporal modelling, and split-plot designs. The class generates relevant matrices and statistics and a wide range of methods including full likelihood estimation of generalised linear mixed models using stochastic Maximum Likelihood, Laplace approximation, power calculation, and access to relevant calculations. The class also includes Hamiltonian Monte Carlo simulation of random effects, sparse matrix methods, and other functionality to support efficient estimation. The \pkg{glmmrOptim} package implements a set of algorithms to identify c-optimal experimental designs where observations are correlated and can be specified using the generalised linear mixed model classes. Several examples and comparisons to existing packages are provided to illustrate use of the packages.

stat.CO

Optimal Study Designs for Cluster Randomised Trials: An Overview of Methods and Results

There are multiple cluster randomised trial designs that vary in when the clusters cross between control and intervention states, when observations are made within clusters, and how many observations are made at that time point. Identifying the most efficient study design is complex though, owing to the correlation between observations within clusters and over time. In this article, we present a review of statistical and computational methods for identifying optimal cluster randomised trial designs. We also adapt methods from the experimental design literature for experimental designs with correlated observations to the cluster trial context. We identify three broad classes of methods: using exact formulae for the treatment effect estimator variance for specific models to derive algorithms or weights for cluster sequences; generalised methods for estimating weights for experimental units; and, combinatorial optimisation algorithms to select an optimal subset of experimental units. We also discuss methods for rounding weights to whole numbers of clusters and extensions to non-Gaussian models. We present results from multiple cluster trial examples that compare the different methods, including problems involving determining optimal allocation of clusters across a set of cluster sequences, and selecting the optimal number of single observations to make in each cluster-period for both Gaussian and non-Gaussian models, and including exchangeable and exponential decay covariance structures.

stat.ME

Efficient design of geographically-defined clusters with spatial autocorrelation

Clusters form the basis of a number of research study designs including survey and experimental studies. Cluster-based designs can be less costly but also less efficient than individual-based designs due to correlation between individuals within the same cluster. Their design typically relies on \textit{ad hoc} choices of correlation parameters, and is insensitive to variations in cluster design. This article examines how to efficiently design clusters where they are geographically defined by demarcating areas incorporating individuals and households or other units. Using geostatistical models for spatial autocorrelation we generate approximations to within cluster average covariance in order to estimate the effective sample size given particular cluster design parameters. We show how the number of enumerated locations, cluster area, proportion sampled, and sampling method affect the efficiency of the design and consider the optimization problem of choosing the most efficient design subject to budgetary constraints. We also consider how the parameters from these approximations can be interpreted simply in terms of `real-world' quantities and used in design analysis.

stat.ME