SearcharxivSearch

arXiv subjects

Jonathan Rougier

Publications and source records attributed to Jonathan Rougier.

12 recordsLinked to original sources

The Double Emulator

Computer models (simulators) are vital tools for investigating physical processes. Despite their utility, the prohibitive run-time of simulators hinders their direct application for uncertainty quantification. Gaussian process emulators (GPEs) have been used extensively to circumvent the cost of the simulator and are known to perform well on simulators with smooth, stationary output. In reality, many simulators violate these assumptions. Motivated by a finite element simulator which models early stage corrosion of uranium in water vapor, we propose an adaption of the GPE, called the double emulator, specifically for simulators which 'ground' in a considerable volume of their input space. Grounding is the process by which a simulator attains its minimum and can result in violation of the stationarity and smoothness assumptions used in the conventional GPE. We perform numerical experiments comparing the performance of the GPE and double emulator on both the corrosion simulator and synthetic examples.

stat.CO

The exact form of the 'Ockham factor' in model selection

We explore the arguments for maximizing the `evidence' as an algorithm for model selection. We show, using a new definition of model complexity which we term `flexibility', that maximizing the evidence should appeal to both Bayesian and Frequentist statisticians. This is due to flexibility's unique position in the exact decomposition of log-evidence into log-fit minus flexibility. In the Gaussian linear model, flexibility is asymptotically equal to the Bayesian Information Criterion (BIC) penalty, but we caution against using BIC in place of flexibility for model selection.

math.ST

Multi-Scale Process Modelling and Distributed Computation for Spatial Data

Recent years have seen a huge development in spatial modelling and prediction methodology, driven by the increased availability of remote-sensing data and the reduced cost of distributed-processing technology. It is well known that modelling and prediction using infinite-dimensional process models is not possible with large data sets, and that both approximate models and, often, approximate-inference methods, are needed. The problem of fitting simple global spatial models to large data sets has been solved through the likes of multi-resolution approximations and nearest-neighbour techniques. Here we tackle the next challenge, that of fitting complex, nonstationary, multi-scale models to large data sets. We propose doing this through the use of superpositions of spatial processes with increasing spatial scale and increasing degrees of nonstationarity. Computation is facilitated through the use of Gaussian Markov random fields and parallel Markov chain Monte Carlo based on graph colouring. The resulting model allows for both distributed computing and distributed data. Importantly, it provides opportunities for genuine model and data scaleability and yet is still able to borrow strength across large spatial scales. We illustrate a two-scale version on a data set of sea-surface temperature containing on the order of one million observations, and compare our approach to state-of-the-art spatial modelling and prediction methods.

stat.CO

Bayesian model-data synthesis with an application to global Glacio-Isostatic Adjustment

We introduce a framework for updating large scale geospatial processes using a model-data synthesis method based on Bayesian hierarchical modelling. Two major challenges come from updating large-scale Gaussian process and modelling non-stationarity. To address the first, we adopt the SPDE approach that uses a sparse Gaussian Markov random fields (GMRF) approximation to reduce the computational cost and implement the Bayesian inference by using the INLA method. For non-stationary global processes, we propose two general models that accommodate commonly-seen geospatial problems. Finally, we show an example of updating an estimate of global glacial isostatic adjustment (GIA) using GPS measurements.

stat.AP

A sparse linear algebra algorithm for fast computation of prediction variances with Gaussian Markov random fields

Gaussian Markov random fields are used in a large number of disciplines in machine vision and spatial statistics. The models take advantage of sparsity in matrices introduced through the Markov assumptions, and all operations in inference and prediction use sparse linear algebra operations that scale well with dimensionality. Yet, for very high-dimensional models, exact computation of predictive variances of linear combinations of variables is generally computationally prohibitive, and approximate methods (generally interpolation or conditional simulation) are typically used instead. A set of conditions are established under which the variances of linear combinations of random variables can be computed exactly using the Takahashi recursions. The ensuing computational simplification has wide applicability and may be used to enhance several software packages where model fitting is seated in a maximum-likelihood framework. The resulting algorithm is ideal for use in a variety of spatial statistical applications, including \emph{LatticeKrig} modelling, statistical downscaling, and fixed rank kriging. It can compute hundreds of thousands exact predictive variances of linear combinations on a standard desktop with ease, even when large spatial GMRF models are used.

stat.CO

A representation theorem for stochastic processes with separable covariance functions, and its implications for emulation

Many applications require stochastic processes specified on two- or higher-dimensional domains; spatial or spatial-temporal modelling, for example. In these applications it is attractive, for conceptual simplicity and computational tractability, to propose a covariance function that is separable; e.g., the product of a covariance function in space and one in time. This paper presents a representation theorem for such a proposal, and shows that all processes with continuous separable covariance functions are second-order identical to the product of second-order uncorrelated processes. It discusses the implications of separable or nearly separable prior covariances for the statistical emulation of complicated functions such as computer codes, and critically reexamines the conventional wisdom concerning emulator structure, and size of design.

math.ST

Exchangeability, the 'Histogram Theorem', and population inference

Some practical results are derived for population inference based on a sample, under the two qualitative conditions of 'ignorability' and exchangeability. These are the 'Histogram Theorem', for predicting the outcome of a non-sampled member of the population, and its application to inference about the population, both without and with groups. There are discussions of parametric versus non-parametric models, and different approaches to marginalisation. An Appendix gives a self-contained proof of the Representation Theorem for finite exchangeable sequences.

math.ST

Rapidly bounding the exceedance probabilities of high aggregate losses

We consider the task of assessing the righthand tail of an insurer's loss distribution for some specified period, such as a year. We present and analyse six different approaches: four upper bounds, and two approximations. We examine these approaches under a variety of conditions, using a large event loss table for US hurricanes. For its combination of tightness and computational speed, we favour the Moment bound. We also consider the appropriate size of Monte Carlo simulations, and the imposition of a cap on single event losses. We strongly favour the Gamma distribution as a flexible model for single event losses, for its tractable form in all of the methods we analyse, its generalisability, and because of the ease with which a cap on losses can be incorporated.

stat.AP

Uncertainty in climate science and climate policy

This essay, written by a statistician and a climate scientist, describes our view of the gap that exists between current practice in mainstream climate science, and the practical needs of policymakers charged with exploring possible interventions in the context of climate change. By `mainstream' we mean the type of climate science that dominates in universities and research centres, which we will term `academic' climate science, in contrast to `policy' climate science; aspects of this distinction will become clearer in what follows. In a nutshell, we do not think that academic climate science equips climate scientists to be as helpful as they might be, when involved in climate policy assessment. Partly, we attribute this to an over-investment in high resolution climate simulators, and partly to a culture that is uncomfortable with the inherently subjective nature of climate uncertainty.

physics.soc-ph

Computation and Visualisation for large-scale Gaussian updates

In geostatistics, and also in other applications in science and engineering, we are now performing updates on Gaussian process models with many thousands or even millions of components. These large-scale inferences involve computational challenges, because the updating equations cannot be solved as written, owing to the size and cost of the matrix operations. They also involve representational challenges, to account for judgements of heterogeneity concerning the underlying fields, and diverse sources of observations. Diagnostics are particularly valuable in this situation. We present a diagnostic and visualisation tool for large-scale Gaussian updates, the `medal plot'. This shows the updated uncertainty for each observation, and also summarises the sharing of information across observations, as a proxy for the sharing of information across the state vector. It allows us to `sanity-check' the code implementing the update, but it can also reveal unexpected features in our modelling. We discuss computational issues for large-scale updates, and we illustrate with an application to assess mass trends in the Antarctic Ice Sheet.

stat.CO

On the use of simple dynamical systems for climate predictions: A Bayesian prediction of the next glacial inception

Over the last few decades, climate scientists have devoted much effort to the development of large numerical models of the atmosphere and the ocean. While there is no question that such models provide important and useful information on complicated aspects of atmosphere and ocean dynamics, skillful prediction also requires a phenomenological approach, particularly for very slow processes, such as glacial-interglacial cycles. Phenomenological models are often represented as low-order dynamical systems. These are tractable, and a rich source of insights about climate dynamics, but they also ignore large bodies of information on the climate system, and their parameters are generally not operationally defined. Consequently, if they are to be used to predict actual climate system behaviour, then we must take very careful account of the uncertainty introduced by their limitations. In this paper we consider the problem of the timing of the next glacial inception, about which there is on-going debate. Our model is the three-dimensional stochastic system of Saltzman and Maasch (1991), and our inference takes place within a Bayesian framework that allows both for the limitations of the model as a description of the propagation of the climate state vector, and for parametric uncertainty. Our inference takes the form of a data assimilation with unknown static parameters, which we perform with a variant on a Sequential Monte Carlo technique (`particle filter'). Provisional results indicate peak glacial conditions in 60,000 years.

physics.ao-ph