SearcharxivSearch

arXiv subjects

Vinny Davies

Publications and source records attributed to Vinny Davies.

9 recordsLinked to original sources

A Bayesian Bi-Directional Splitting Framework for Variable Selection in Large Datasets

Modern tabular datasets are becoming increasingly large, both in the number of samples and covariates, posing significant challenges for Bayesian variable selection due to the resulting computational burden. While there is extensive literature on scaling Bayesian inference to large numbers of observations or high-dimensional covariate spaces, comparatively little work addresses scalable Bayesian variable selection when both dimensions are large simultaneously in a practical setting. This paper presents a novel Bayesian variable selection framework for efficiently analysing data with a large number of both rows and columns. The proposed framework operates via a divide-and-conquer approach, splitting data into batches along both directions and analysing each batch independently in parallel, after which the results are combined together in a two-phase consensus procedure. Experiments demonstrate the computational gains while retaining strong variable selection performance, successfully identifying relevant signals in noisy, high-dimensional settings despite the loss of information induced by data splitting. Practical guidelines are also provided, including recommendations for tuning parameter choices and effective data partitioning strategies, followed by a real application to an H3N2 influenza dataset.

stat.ME

Emulation strategies for Bayesian inference of regional left ventricle material parameters

Patient-specific biomechanical models of the left ventricle can relate cardiac magnetic resonance imaging to regional myocardial material properties, but existing emulator-based studies typically treat the myocardium as mechanically homogeneous, limiting representation of localised dysfunction. We propose a Bayesian surrogate-modelling framework for inferring regional Holzapfel-Ogden material parameters in a left ventricle partitioned into five physiological zones derived from the American Heart Association 17-segment model. Eight emulator strategies spanning single- versus multi-output, local versus global, and Gaussian-process- versus neural-network-based architectures were screened using parameter point-estimation accuracy; the three retained models were evaluated using empirical marginal credible-interval coverage and posterior contraction. We found that models with comparable point accuracy nevertheless differed markedly in uncertainty. A multi-output variational Gaussian process provided the most favourable balance across these criteria and was retained for the subsequent analyses. In synthetic local and global stiffening scenarios, maximum a posteriori estimates generally distinguished stiffened from baseline zones, but the nonlinear-stiffening parameters were more difficult to identify from end-diastolic observations than the stiffness-magnitude parameters. A healthy-volunteer analysis demonstrates feasibility with an incomplete observation vector and jointly inferred strain-noise scales. These results suggest that the proposed framework provides a computationally feasible, uncertainty-aware approach to regional left ventricle parameter inference.

stat.AP

Spatial prediction of environmental processes using random forests: How best to account for spatial dependence?

Geostatistical spatial prediction for environmental processes is typically undertaken using Gaussian process models via Kriging, while machine learning (ML) algorithms are state-of-the-art for non-spatial prediction. An exciting recent fusion of these ideas imbibes traditional ML algorithms with the capacity to deal with spatial autocorrelation, leading to improved predictive performance. A range of approaches have been proposed, including fusion with Gaussian processes, observation-driven correlation structures, spatial basis functions and local geographical fitting. However, there has been no numerical comparison of their relative predictive performances, which is needed to advise environmental scientists on the optimal approach to use. This paper fills this knowledge gap, and focuses on random forests as the ML algorithm because they are more computationally and conceptually straightforward to implement than deep learning algorithms. The results from two studies are presented, the first being a controlled simulation experiment investigating whether any single approach is consistently superior across different spatial autocorrelation types. The second study focuses on the prediction of air pollution concentrations within a tuberculosis prevalence study in Blantyre, Malawi. The results show that whilst no single approach is universally superior, utilising spatial basis functions appears to perform consistently well across both the simulation and real data studies.

stat.ME

Reflections on the Future of Statistics Education in a Technological Era

Keeping pace with rapidly evolving technology is a key challenge in teaching statistics. To equip students with essential skills for the modern workplace, educators must integrate relevant technologies into the statistical curriculum where possible. University-level statistics education has experienced substantial technological change, particularly in the tools and practices that underpin teaching and learning. Statistical programming has become central to many courses, with R widely used and Python increasingly incorporated into statistics and data analytics programmes. Additionally, coding practices, database management, and machine learning now feature within some statistics curricula. Looking ahead, we anticipate a growing emphasis on artificial intelligence (AI), particularly the pedagogical implications of generative AI tools such as ChatGPT. In this article, we explore these technological developments and discuss strategies for their integration into contemporary statistics education.

stat.OT

A spatial random forest algorithm for population-level epidemiological risk assessment

Spatial epidemiology identifies the drivers of elevated population-level disease risks, using disease counts, exposures and known confounders at the areal unit level. Poisson regression models are typically used for inference, which incorporate a linear/additive regression component and allow for unmeasured confounding via a set of spatially autocorrelated random effects. This approach requires the confounder interactions and their functional relationships with disease risk to be specified in advance, rather than being learned from the data. Therefore, this paper proposes the SPAR-Forest-ERF algorithm, which is the first fusion of random forests for capturing non-linear and interacting confounder-response effects with Bayesian spatial autocorrelation models that can estimate interpretable exposure response functions (ERF) with full uncertainty quantification. Methodologically, we extend existing methods set in a prediction context by propagating uncertainty between both the ML and statistical models, developing a new stopping criteria designed to ensure the stability of the primary inferential target, and incorporating a range of different ERFs for maximum model flexibility. This methodology is motivated by a new study quantifying the impact of air pollution concentrations on self-rated health in Scotland, using data from the recently released 2022 national census.

stat.ME

Conditional autoregressive models fused with random forests to improve small-area spatial prediction

In areal unit data with missing or suppressed data, it desirable to create models that are able to predict observations that are not available. Traditional statistical methods achieve this through Bayesian hierarchical models that can capture the unexplained residual spatial autocorrelation through conditional autoregressive (CAR) priors, such that they can make predictions at geographically related spatial locations. In contrast, typical machine learning approaches such as random forests ignore this residual autocorrelation, and instead base predictions on complex non-linear feature-target relationships. In this paper, we propose CAR-Forest, a novel spatial prediction algorithm that combines the best features of both approaches by fusing them together. By iteratively refitting a random forest combined with a Bayesian CAR model in one algorithm, CAR-Forest can incorporate flexible feature-target relationships while still accounting for the residual spatial autocorrelation. Our results, based on a Scottish housing price data set, show that CAR-Forest outperforms Bayesian CAR models, random forests, and the state-of-the-art hybrid approach, geographically weighted random forest, providing a state-of-the-art framework for small-area spatial prediction.

stat.ME

Generalised linear models for prognosis and intervention: Theory, practice, and implications for machine learning

Prediction and causal explanation are fundamentally distinct tasks of data analysis. In health applications, this difference can be understood in terms of the difference between prognosis (prediction) and prevention/treatment (causal explanation). Nevertheless, these two concepts are often conflated in practice. We use the framework of generalised linear models (GLMs) to illustrate that predictive and causal queries require distinct processes for their application and subsequent interpretation of results. In particular, we identify five primary ways in which GLMs for prediction differ from GLMs for causal inference: (1) The covariates that should be considered for inclusion in (and possibly exclusion from) the model; (2) How a suitable set of covariates to include in the model is determined; (3) Which covariates are ultimately selected, and what functional form (i.e. parameterisation) they take; (4) How the model is evaluated; and (5) How the model is interpreted. We outline some of the potential consequences of failing to acknowledge and respect these differences, and additionally consider the implications for machine learning (ML) methods. We then conclude with three recommendations which we hope will help ensure that both prediction and causal modelling are used appropriately and to greatest effect in health research.

stat.AP

Fast Parameter Inference in a Biomechanical Model of the Left Ventricle using Statistical Emulation

A central problem in biomechanical studies of personalised human left ventricular (LV) modelling is estimating the material properties and biophysical parameters from in-vivo clinical measurements in a time frame suitable for use within a clinic. Understanding these properties can provide insight into heart function or dysfunction and help inform personalised medicine. However, finding a solution to the differential equations which mathematically describe the kinematics and dynamics of the myocardium through numerical integration can be computationally expensive. To circumvent this issue, we use the concept of emulation to infer the myocardial properties of a healthy volunteer in a viable clinical time frame using in-vivo magnetic resonance image (MRI) data. Emulation methods avoid computationally expensive simulations from the LV model by replacing the biomechanical model, which is defined in terms of explicit partial differential equations, with a surrogate model inferred from simulations generated before the arrival of a patient, vastly improving computational efficiency at the clinic. We compare and contrast two emulation strategies: (i) emulation of the computational model outputs and (ii) emulation of the loss between the observed patient data and the computational model outputs. These strategies are tested with two different interpolation methods, as well as two different loss functions...

stat.AP

Improving the identification of antigenic sites in the H1N1 Influenza virus through accounting for the experimental structure in a sparse hierarchical Bayesian model

Understanding how genetic changes allow emerging virus strains to escape the protection afforded by vaccination is vital for the maintenance of effective vaccines. In the current work, we use structural and phylogenetic differences between pairs of virus strains to identify important antigenic sites on the surface of the influenza A(H1N1) virus through the prediction of haemagglutination inhibition (HI) assay, pairwise measures of the antigenic similarity of virus strains. We propose a sparse hierarchical Bayesian model that can deal with the pairwise structure and inherent experimental variability in the H1N1 data through the introduction of latent variables. The latent variables represent the underlying HI assay measurement of any given pair of virus strains and help account for the fact that for any HI assay measurement between the same pair of virus strains, the difference in the viral sequence remains the same. Through accurately representing the structure of the H1N1 data, the model is able to select virus sites which are antigenic, while its latent structure achieves the computational efficiency required to deal with large virus sequence data, as typically available for the influenza virus. In addition to the latent variable model, we also propose a new method, block integrated Widely Applicable Information Criterion (biWAIC), for selecting between competing models. We show how this allows us to effectively select the random effects when used with the proposed model and apply both methods to an A(H1N1) dataset.

stat.AP