SearcharxivSearch

arXiv subjects

Daniel Cooley

Publications and source records attributed to Daniel Cooley.

16 recordsLinked to original sources

A Vector Space Approach to Heavy Tailed Analysis

We construct a vector space whose defining characteristics are rooted in univariate regular variation of random variables. Specifically, the base vector space $\mathbb{V}_b$ consists of random variables whose limiting tail probabilities, when scaled by regularly varying functions of the form $b(s)=s^\alpha L(s)$, are finite. Defining a subspace ${\cal N}_b$ corresponding to random variables in $\mathbb{V}_b$ whose limiting tail probabilities are zero when normalized by $b(s)$ allows the base space $\mathbb{V}_b$ to be partitioned into equivalence classes. We define a vector space $\mathbb{W}_b$ consisting of these equivalence classes, and show its nonzero elements are equivalence classes of regularly varying random variables. We show that a natural norm exists for $\mathbb{W}_b$ if $\alpha > 1$. We show that the equivalence classes and convergence in norm are different than more familiar vector spaces of random variables. Turning our attention to extreme value modeling, we consider finite-dimensional subspaces of $\mathbb{W}_b$ whose basis vectors are jointly regularly varying. We show that in the case $\alpha = 2$, the previously defined tail pairwise dependence measure serves as an inner product. As any finite-dimensional space is complete, we can use the projection theorem to perform linear prediction.

math.PR

Quantifying Very Extreme Precipitation and Temperature Using Huge Ensembles Generated by Machine Learning-based Climate Model Emulators

Weather extremes produce major impacts on society and ecosystems and are likely to change in likelihood and magnitude with climate change. However, very low probability events are hard to characterize statistically using observations or even climate model output because of short records/runs. For precipitation, consideration of such events arises in quantifying Probable Maximum Precipitation (PMP), namely estimating extreme precipitation magnitudes for designing and assessing critical infrastructure. A recent National Academies report on modernizing PMP estimation proposed using very large climate model-based ensembles to estimate extreme quantiles, possibly through machine learning-based ensemble boosting. Here we assess statistical aspects of such an approach for the contiguous United States using a huge ensemble (10560 years) produced by a state-of-the-art emulator (ACE2) trained on ERA5 reanalysis. The results indicate that one can practically estimate very extreme precipitation and temperature quantiles, provided one uses appropriate statistical extreme value techniques. More specifically, the results provide evidence for (1) the use of threshold-exceedance methods with a sufficiently high threshold (necessary for precipitation) for reliable estimation, (2) the robustness of results to variation in extremes by season and storm type, and (3) the sufficiency of the ensemble for well-constrained statistical uncertainty. Our results also show that the emulator produces extremes outside the range of the ERA5 training data. While encouraging for emulators' potential use for quantifying the climatology of extremes, more investigation is needed to assess whether emulators are fit for this purpose. Our focus is on how to use huge ensembles to estimate very extreme statistics; we expect the results to be relevant for future improved emulators.

stat.AP

Transformed-Linear Innovations Algorithm for Modeling and Forecasting of Time Series Extremes

The innovations algorithm is a classical recursive forecasting algorithm used in time series analysis. We develop the innovations algorithm for a class of nonnegative regularly varying time series models constructed via transformed-linear arithmetic. In addition to providing the best linear predictor, the algorithm also enables us to estimate parameters of transformed-linear regularly-varying moving average (MA) models, thus providing a tool for modeling. We first construct an inner product space of transformed-linear combinations of nonnegative regularly-varying random variables and prove its link to a Hilbert space which allows us to employ the projection theorem, from which we develop the transformed-linear innovations algorithm. Turning our attention to the class of transformed linear MA($\infty$) models, we give results on parameter estimation and also show that this class of models is dense in the class of possible tail pairwise dependence functions (TPDFs). We also develop an extremes analogue of the classical Wold decomposition. Simulation study shows that our class of models captures tail dependence for the GARCH(1,1) model and a Markov time series model, both of which are outside our class of models.

math.ST

Semiparametric Estimation of the Shape of the Limiting Bivariate Point Cloud

We propose a model to flexibly estimate joint tail properties by exploiting the convergence of an appropriately scaled point cloud onto a compact limit set. Characteristics of the shape of the limit set correspond to key tail dependence properties. We directly model the shape of the limit set using Bezier splines, which allow flexible and parsimonious specification of shapes in two dimensions. We fit the Bezier splines to data in pseudo-polar coordinates using Markov chain Monte Carlo sampling, utilizing a limiting approximation to the conditional likelihood of the radii given angles. We propose a novel prior on the shape of the limit set via constraints on the parameters of the Bezier splines. A direct advantage of our Bayesian approach is that the support of this prior guarantees that each posterior sample is a valid limit set boundary, allowing direct posterior analysis of any quantity derived from the shape of the curve. Furthermore, we obtain interpretable inference on the asymptotic dependence class by using mixture priors with point masses on the corner of the unit box. Finally, we apply our model to bivariate datasets of extremes of variables related to fire risk and air pollution.

stat.ME

Autonomous Apple Fruitlet Sizing and Growth Rate Tracking using Computer Vision

In this paper, we present a computer vision-based approach to measure the sizes and growth rates of apple fruitlets. Measuring the growth rates of apple fruitlets is important because it allows apple growers to determine when to apply chemical thinners to their crops in order to optimize yield. The current practice of obtaining growth rates involves using calipers to record sizes of fruitlets across multiple days. Due to the number of fruitlets needed to be sized, this method is laborious, time-consuming, and prone to human error. With images collected by a hand-held stereo camera, our system, segments, clusters, and fits ellipses to fruitlets to measure their diameters. The growth rates are then calculated by temporally associating clustered fruitlets across days. We provide quantitative results on data collected in an apple orchard, and demonstrate that our system is able to predict abscise rates within 3.5% of the current method with a 6 times improvement in speed, while requiring significantly less manual effort. Moreover, we provide results on images captured by a robotic system in the field, and discuss the next steps required to make the process fully autonomous.

cs.RO

Transformed Linear Prediction for Extremes

We address the problem of prediction for extreme observations by proposing an extremal linear prediction method. We construct an inner product space of nonnegative random variables derived from transformed-linear combinations of independent regularly varying random variables. Under a reasonable modeling assumption, the matrix of inner products corresponds to the tail pairwise dependence matrix, which can be easily estimated. We derive the optimal transformed-linear predictor via the projection theorem, which yields a predictor with the same form as the best linear unbiased predictor in non-extreme settings. We quantify uncertainty for prediction errors by constructing prediction intervals based on the geometry of regular variation. We demonstrate the effectiveness of our method through a simulation study and its applications to predicting high pollution levels, and extreme precipitation.

stat.ME

Simulating flood event sets using extremal principal components

Hazard event sets, a collection of synthetic extreme events over a given period, are important for catastrophe modelling. This paper addresses the issue of generating event sets of extreme river flow for northern England and southern Scotland, a region which has been particularly affected by severe flooding over the past 20 years. We start by analysing historical extreme river flow across 45 gauges, located within the study region, using methods from extreme value analysis, including the concept of extremal principal components. Our analysis reveals interesting connections between the extremal dependence structure and the region's topography/climate. We then introduce a framework which is based on modelling the distribution of the extremal principal components in order to generate synthetic events of extreme river flow. The generative framework is dimension-reducing in that it distinctly handles the principal components based on their contribution to describing the nature of extreme river flow across the study region. We also detail a data-driven approach to select the optimal dimension. Synthetic flood events are subsequently generated efficiently by sampling from the fitted distribution. Our approach for generating hazard event sets can be easily implemented by practitioners and our results indicate good agreement between the observed and simulated extreme river flow dynamics. For the considered application, we also find that our approach outperforms existing statistical approaches for generating hazard event sets.

stat.AP

Transformed-Linear Models for Time Series Extremes

In order to capture the dependence in the upper tail of a time series, we develop non-negative regularly-varying time series models that are constructed similarly to classical non-extreme ARMA models. Rather than fully characterizing tail dependence of the time series, we define the concept of weak tail stationarity which allows us to describe a regularly-varying time series through the tail pairwise dependence function (TPDF) which is a measure of pairwise extremal dependencies. We state consistency requirements among the finite-dimensional collections of the elements of a regularly-varying time series and show that the TPDF's value does not depend on the dimension being considered. So that our models take nonnegative values, we use transformed-linear operations. We show existence and stationarity of these models, and develop their properties such as the model TPDF's. Additionally, we show the class of transformed-linear MA($\infty$) models forms an inner product space. Motivated by investigating conditions conducive to the spread of wildfires, we fit models to hourly windspeed data and find that the fitted transformed-linear models produce better estimates of upper tail quantities than traditional ARMA models or than classical linear regularly varying models.

stat.ME

Decompositions of Dependence for High-Dimensional Extremes

Employing the framework of regular variation, we propose two decompositions which help to summarize and describel high-dimensional tail dependence. Via transformation, we define a vector space on the positive orthant, yielding the notion of basis. With a suitably-chosen transformation, we show that transformed-linear operations applied to regularly varying random vectors preserve regular variation. Rather than model regular-variation's angular measure, we summarize tail dependence via a matrix of pairwise tail dependence metrics. This matrix is positive semidefinite, and eigendecomposition allows one to interpret tail dependence via the resulting eigenbasis. Additionally this matrix is completely positive, and a resulting decomposition allows one to easily construct regularly varying random vectors which share the same pairwise tail dependencies. We illustrate our methods with Swiss rainfall data and financial return data.

stat.ME

A Nonparametric Method for Producing Isolines of Bivariate Exceedance Probabilities

We present a method for drawing isolines indicating regions of equal joint exceedance probability for bivariate data. The method relies on bivariate regular variation, a dependence framework widely used for extremes. This framework enables drawing isolines corresponding to very low exceedance probabilities and these lines may lie beyond the range of the data. The method we utilize for characterizing dependence in the tail is largely nonparametric. Furthermore, we extend this method to the case of asymptotic independence and propose a procedure which smooths the transition from asymptotic independence in the interior to the first-order behavior on the axes. We propose a diagnostic plot for assessing isoline estimate and choice of smoothing, and a bootstrap procedure to visually assess uncertainty.

stat.ME

A Markov-switching model for heat waves

Heat waves merit careful study because they inflict severe economic and societal damage. We use an intuitive, informal working definition of a heat wave-a persistent event in the tail of the temperature distribution-to motivate an interpretable latent state extreme value model. A latent variable with dependence in time indicates membership in the heat wave state. The strength of the temporal dependence of the latent variable controls the frequency and persistence of heat waves. Within each heat wave, temperatures are modeled using extreme value distributions, with extremal dependence across time accomplished through an extreme value Markov model. One important virtue of interpretability is that model parameters directly translate into quantities of interest for risk management, so that questions like whether heat waves are becoming longer, more severe or more frequent are easily answered by querying an appropriate fitted model. We demonstrate the latent state model on two recent, calamitous, examples: the European heat wave of 2003 and the Russian heat wave of 2010.

stat.ME

Data Mining to Investigate the Meteorological Drivers for Extreme Ground Level Ozone Events

This project aims to explore which combinations of meteorological conditions are associated with extreme ground level ozone conditions. Our approach focuses only on the tail by optimizing the tail dependence between the ozone response and functions of meteorological covariates. Since there is a long list of possible meteorological covariates, the space of possible models cannot be explored completely. Consequently, we perform data mining within the model selection context, employing an automated model search procedure. Our study is unique among extremes applications as optimizing tail dependence has not previously been attempted, and it presents new challenges, such as requiring a smooth threshold. We present a simulation study which shows that the method can detect complicated conditions leading to extreme responses and resists overfitting. We apply the method to ozone data for Atlanta and Charlotte and find similar meteorological drivers for these two Southeastern US cities. We identify several covariates which help to differentiate the meteorological conditions which lead to extreme ozone levels from those which lead to merely high levels.

stat.AP

Extreme value analysis for evaluating ozone control strategies

Tropospheric ozone is one of six criteria pollutants regulated by the US EPA, and has been linked to respiratory and cardiovascular endpoints and adverse effects on vegetation and ecosystems. Regional photochemical models have been developed to study the impacts of emission reductions on ozone levels. The standard approach is to run the deterministic model under new emission levels and attribute the change in ozone concentration to the emission control strategy. However, running the deterministic model requires substantial computing time, and this approach does not provide a measure of uncertainty for the change in ozone levels. Recently, a reduced form model (RFM) has been proposed to approximate the complex model as a simple function of a few relevant inputs. In this paper, we develop a new statistical approach to make full use of the RFM to study the effects of various control strategies on the probability and magnitude of extreme ozone events. We fuse the model output with monitoring data to calibrate the RFM by modeling the conditional distribution of monitoring data given the RFM using a combination of flexible semiparametric quantile regression for the center of the distribution where data are abundant and a parametric extreme value distribution for the tail where data are sparse. Selected parameters in the conditional distribution are allowed to vary by the RFM value and the spatial location. Also, due to the simplicity of the RFM, we are able to embed the RFM in our Bayesian hierarchical framework to obtain a full posterior for the model input parameters, and propagate this uncertainty to the estimation of the effects of the control strategies. We use the new framework to evaluate three potential control strategies, and find that reducing mobile-source emissions has a larger impact than reducing point-source emissions or a combination of several emission sources.

stat.AP

Approximating the conditional density given large observed values via a multivariate extremes framework, with application to environmental data

Phenomena such as air pollution levels are of greatest interest when observations are large, but standard prediction methods are not specifically designed for large observations. We propose a method, rooted in extreme value theory, which approximates the conditional distribution of an unobserved component of a random vector given large observed values. Specifically, for $\mathbf{Z}=(Z_1,...,Z_d)^T$ and $\mathbf{Z}_{-d}=(Z_1,...,Z_{d-1})^T$, the method approximates the conditional distribution of $[Z_d|\mathbf{Z}_{-d}=\mathbf{z}_{-d}]$ when $|\mathbf{z}_{-d}|>r_*$. The approach is based on the assumption that $\mathbf{Z}$ is a multivariate regularly varying random vector of dimension $d$. The conditional distribution approximation relies on knowledge of the angular measure of $\mathbf{Z}$, which provides explicit structure for dependence in the distribution's tail. As the method produces a predictive distribution rather than just a point predictor, one can answer any question posed about the quantity being predicted, and, in particular, one can assess how well the extreme behavior is represented. Using a fitted model for the angular measure, we apply our method to nitrogen dioxide measurements in metropolitan Washington DC. We obtain a predictive distribution for the air pollutant at a location given the air pollutant's measurements at four nearby locations and given that the norm of the vector of the observed measurements is large.

stat.AP

Bayesian Inference from Composite Likelihoods, with an Application to Spatial Extremes

Composite likelihoods are increasingly used in applications where the full likelihood is analytically unknown or computationally prohibitive. Although the maximum composite likelihood estimator has frequentist properties akin to those of the usual maximum likelihood estimator, Bayesian inference based on composite likelihoods has yet to be explored. In this paper we investigate the use of the Metropolis--Hastings algorithm to compute a pseudo-posterior distribution based on the composite likelihood. Two methodologies for adjusting the algorithm are presented and their performance on approximating the true posterior distribution is investigated using simulated data sets and real data on spatial extremes of rainfall.

stat.ME

Downscaling extremes: A comparison of extreme value distributions in point-source and gridded precipitation data

There is substantial empirical and climatological evidence that precipitation extremes have become more extreme during the twentieth century, and that this trend is likely to continue as global warming becomes more intense. However, understanding these issues is limited by a fundamental issue of spatial scaling: most evidence of past trends comes from rain gauge data, whereas trends into the future are produced by climate models, which rely on gridded aggregates. To study this further, we fit the Generalized Extreme Value (GEV) distribution to the right tail of the distribution of both rain gauge and gridded events. The results of this modeling exercise confirm that return values computed from rain gauge data are typically higher than those computed from gridded data; however, the size of the difference is somewhat surprising, with the rain gauge data exhibiting return values sometimes two or three times that of the gridded data. The main contribution of this paper is the development of a family of regression relationships between the two sets of return values that also take spatial variations into account. Based on these results, we now believe it is possible to project future changes in precipitation extremes at the point-location level based on results from climate models.

stat.AP