SearcharxivSearch

arXiv subjects

Thomas Opitz

Publications and source records attributed to Thomas Opitz.

At least 19 recordsLinked to original sources

Extrapolation of extreme covariates in generalized additive regression using extreme-value theory

We propose methods to enhance the predictive performance of generalized additive models (GAMs) in the context of covariate extrapolation, where predictions rely on covariates beyond their observed range. When using predictive models such as GAMs, shifts in the covariate distribution between training and prediction datasets can occur. Ignoring this issue may lead to inaccurate predictions in the tail of the covariate distributions. For example, this problem is particularly critical in climate-change scenarios, where covariates simulated from future climate scenarios are likely to contain more extreme conditions. Our approach integrates GAMs for the bulk of covariate distributions with asymptotic models from multivariate extreme-value theory at high covariate values. We consider binary responses based on a latent variable assumption, and also continuous responses. For large values of the covariates, on a specific marginal scale motivated by extreme-value theory the latent variable or continuous response is assumed to depend linearly on the covariates with an additive error term, when using an appropriate link function. In an application to wildfires in Europe, we explore how the new method can improve predictions, using environmental and meteorological covariates.

stat.ME

Spatio-temporal modeling of urban extreme rainfall events at high resolution

Modeling precipitation and its accumulation over time and space is essential for flood risk assessment. In this paper, we analyze rainfall data collected over several years through a micro-scale precipitation sensor network in Montpellier, France. A novel spatio-temporal stochastic model is proposed for high-resolution urban extreme rainfall and combines realistic marginal behaviour and flexible dependence structure. Marginally, rainfall intensities are described by the Extended Generalized Pareto Distribution (EGPD), capturing both moderate and extreme events without threshold selection. Based on peaks-over-threshold theory for spatial processes, dependence during extreme episodes is modeled by an r-Pareto process with a non-separable variogram allowing for episode-specific advection, such that the displacement of rainfall cells is represented explicitly. Based on a catalog of extreme space-time episodes extracted from observations, parameters are estimated by a new composite likelihood based on joint exceedance indicators. Empirical advection velocities are derived beforehand from a radar reanalysis dataset. We show that the model accurately reproduces the spatio-temporal structure of extreme rainfall observed in the Montpellier OMSEV network and enables realistic stochastic scenario generation for flood risk assessment.

stat.AP

Continuous mixtures of Gaussian processes as models for spatial extremes

Spatial modelling of extreme values allows studying the risk of joint occurrence of extreme events at different locations and is of significant interest in climatic and other environmental sciences. A popular class of dependence models for spatial extremes is that of random location-scale mixtures, in which a spatial "baseline" process is multiplied or shifted by a random variable, potentially altering its extremal dependence behaviour. Gaussian location-scale mixtures retain benefits of their Gaussian baseline processes while overcoming some of their limitations, such as symmetry, light tails and weak tail dependence. We review properties of Gaussian location-scale mixtures and develop novel constructions with interesting features, together with a general algorithm for conditional simulation from these models. We leverage their flexibility to propose extended extreme-value models, that allow for appropriately modelling not only the tails but also the bulk of the data. This is important in many applications and avoids the need to explicitly select the events considered as extreme. We propose new solutions for likelihood inference in parametric models of Gaussian location-scale mixtures, in order to avoid the numerical bottleneck given by the latent location and scale variables that can lead to high computational cost of standard likelihood evaluations. The effectiveness of the models and of the inference methods is confirmed with simulated data examples, and we present an application to wildfire-related weather variables in Portugal. Although not detailed here, the approaches would also be straightforward to use for modelling multivariate (non spatial) data.

stat.ME

Modeling high and low extremes with a novel dynamic spatio-temporal model

Extreme environmental events such as severe storms, drought, heat waves, flash floods, and abrupt species collapse have become more prevalent in the earth-atmosphere dynamic system in recent years. In order to fully understand the underlying mechanisms and enhance informed decision-making, a flexible model capable of accommodating extremes is necessary. Existing dynamic spatio-temporal statistical models exhibit limitations in capturing extremes when assuming Gaussian error distributions, whereas the current models for spatial extremes mostly assume temporal independence and are focused on joint upper tails at two or more locations. Here, we introduce a new class of dynamic spatio-temporal models that capture both high and low extremes using a mixture of heavy- and light-tailed distributions with varying tail indices. Our framework flexibly identifies extremal dependence and independence in both space and time with uncertainty quantification and supports missing data prediction, as in other dynamic spatio-temporal models. We demonstrate its effectiveness using a large reanalysis dataset of hourly particulate matter in the Central United States.

stat.ME

On the spatial extent of extreme threshold exceedances

We introduce the extremal range, a local statistic for studying the spatial extent of extreme events in random fields on $\mathbb{R}^d$. Conditioned on exceedance of a high threshold at a location $s$, the extremal range at $s$ is the random variable defined as the smallest distance from $s\in\mathbb{R}^d$ to a location where there is a nonexceedance. We leverage tools from excursion-set theory, such as Lipschitz- Killing curvatures, to express distributional properties of the extremal range, including asymptotics for small distances and high thresholds. The extremal range captures the rate at which the spatial extent of conditional extreme events scales for increasingly high thresholds, and we relate its distributional properties with the well-known bivariate tail dependence coefficient and the extremal index of time series in Extreme-Value Theory. We calculate theoretical extremal-range properties for commonly used models, such as Gaussian or regularly varying random fields. Numerical studies illustrate that, when the extremal range is estimated from discretized excursion sets observed on compact observation windows, the distribution of the resulting estimators appropriately reproduces the theoretically derived links with the Lipschitz- Killing curvature densities.

math.ST

Pareto processes for threshold exceedances in spatial extremes

We review some recent development in the theory of spatial extremes related to Pareto Processes and modeling of threshold exceedances. We provide theoretical background, methodology for modeling, simulation and inference as well as an illustration to wave height modelling. This preprint is an author version of a chapter to appear in a collaborative book.

math.ST

Modeling of spatial extremes in environmental data science: Time to move away from max-stable processes

Environmental data science for spatial extremes has traditionally relied heavily on max-stable processes. Even though the popularity of these models has perhaps peaked with statisticians, they are still perceived and considered as the `state-of-the-art' in many applied fields. However, while the asymptotic theory supporting the use of max-stable processes is mathematically rigorous and comprehensive, we think that it has also been overused, if not misused, in environmental applications, to the detriment of more purposeful and meticulously validated models. In this paper, we review the main limitations of max-stable process models, and strongly argue against their systematic use in environmental studies. Alternative solutions based on more flexible frameworks using the exceedances of variables above appropriately chosen high thresholds are discussed, and an outlook on future research is given, highlighting recommendations moving forward and the opportunities offered by hybridizing machine learning with extreme-value statistics.

stat.ME

Extreme-value modelling of migratory bird arrival dates: Insights from citizen science data

Citizen science mobilises many observers and gathers huge datasets but often without strict sampling protocols, resulting in observation biases due to heterogeneous sampling effort, which can lead to biased statistical inferences. We develop a spatiotemporal Bayesian hierarchical model for bias-corrected estimation of arrival dates of the first migratory bird individuals at a breeding site. Higher sampling effort could be correlated with earlier observed dates. We implement data fusion of two citizen-science datasets with fundamentally different protocols (BBS, eBird) and map posterior distributions of the latent process, which contains four spatial components with Gaussian process priors: species niche; sampling effort; position and scale parameters of annual first arrival date. The data layer includes four response variables: counts of observed eBird locations (Poisson); presence-absence at observed eBird locations (Binomial); BBS occurrence counts (Poisson); first arrival dates (Generalised Extreme-Value). We devise a Markov Chain Monte Carlo scheme and check by simulation that the latent process components are identifiable. We apply our model to several migratory bird species in the northeastern US for 2001--2021 and find that the sampling effort significantly modulates the observed first arrival date. We exploit this relationship to effectively bias-correct predictions of the true first arrivals.

stat.ME

Assessing the size of spatial extreme events using local coefficients based on excursion sets

Extreme events arising in georeferenced processes can take various forms, such as occurring in isolated patches or stretching contiguously over large areas, and can further vary with the spatial location and the extremeness of the events. We use excursion sets above threshold exceedances in data observed over a two-dimensional grid of rectangular pixels to propose a general family of coefficients that assess spatial-extent properties relevant for risk assessment, and study five candidate coefficients from this family. These coefficients are defined locally and interpreted as a spatial distance from a reference site where the threshold is exceeded. We develop statistical inference and discuss robustness to boundary effects and resolution of the pixel grid. To statistically extrapolate coefficients towards very high threshold levels, we formulate a semiparametric model and estimate a parameter characterizing how coefficients scale with the quantile level of the threshold. The utility of the new coefficients is illustrated through simulated data, as well as in an application to gridded daily temperature in continental France. We find notable differences in estimated coefficient maps between climate model simulations and observation-based reanalysis.

math.ST

On the perimeter estimation of pixelated excursion sets of 2D anisotropic random fields

We are interested in creating statistical methods to provide informative summaries of random fields through the geometry of their excursion sets. To this end, we introduce an estimator for the length of the perimeter of excursion sets of random fields on $\mathbb{R}^2$ observed over regular square tilings. The proposed estimator acts on the empirically accessible binary digital images of the excursion regions and computes the length of a piecewise linear approximation of the excursion boundary. The estimator is shown to be consistent as the pixel size decreases, without the need of any normalization constant, and with neither assumption of Gaussianity nor isotropy imposed on the underlying random field. In this general framework, even when the domain grows to cover $\mathbb{R}^2$, the estimation error is shown to be of smaller order than the side length of the domain. For affine, strongly mixing random fields, this translates to a multivariate Central Limit Theorem for our estimator when multiple levels are considered simultaneously. Finally, we conduct several numerical studies to investigate statistical properties of the proposed estimator in the finite-sample data setting.

math.ST

Spatial modeling and future projection of extreme precipitation extents

Extreme precipitation events with large spatial extents may have more severe impacts than localized events as they can lead to widespread flooding. It is debated how climate change may affect the spatial extent of precipitation extremes, whose investigation often directly relies on simulations from climate models. Here, we use a different strategy to investigate how future changes in spatial extents of precipitation extremes differ across climate zones and seasons in two river basins (Danube and Mississippi). We rely on observed precipitation extremes while exploiting a physics-based mean temperature covariate, which enables us to project future precipitation extents. We include the covariate into newly developed time-varying $r$-Pareto processes using a suitably chosen spatial aggregation functional $r$. This model captures temporal non-stationarity in the spatial dependence structure of precipitation extremes by linking it to the temperature covariate, which we derive from observations for model calibration and from debiased climate simulations (CMIP6) for projections. For both river basins, our results show negative correlation between the spatial extent and the temperature covariate for most of the rain season and an increasing trend in the margins, indicating a decrease in spatial precipitation extent in a warming climate during rain seasons as precipitation intensity increases locally.

stat.ME

Partial Tail-Correlation Coefficient Applied to Extremal-Network Learning

We propose a novel extremal dependence measure called the partial tail-correlation coefficient (PTCC), in analogy to the partial correlation coefficient in classical multivariate analysis. The construction of our new coefficient is based on the framework of multivariate regular variation and transformed-linear algebra operations. We show how this coefficient allows identifying pairs of variables that have partially uncorrelated tails given the other variables in a random vector. Unlike other recently introduced conditional independence frameworks for extremes, our approach requires minimal modeling assumptions and can thus be used in exploratory analyses to learn the structure of extremal graphical models. Similarly to traditional Gaussian graphical models where edges correspond to the non-zero entries of the precision matrix, we can exploit classical inference methods for high-dimensional data, such as the graphical LASSO with Laplacian spectral constraints, to efficiently learn the extremal network structure via the PTCC. We apply our new method to study extreme risk networks in two different datasets (extreme river discharges and historical global currency exchange data) and show that we can extract meaningful extremal structures with meaningful domain-specific interpretations.

stat.ME

Joint modeling of landslide counts and sizes using spatial marked point processes with sub-asymptotic mark distributions

To accurately quantify landslide hazard in a region of Turkey, we develop new marked point process models within a Bayesian hierarchical framework for the joint prediction of landslide counts and sizes. To accommodate for the dominant role of the few largest landslides in aggregated sizes, we leverage mark distributions with strong justification from extreme-value theory, thus bridging the two broad areas of statistics of extremes and marked point patterns. At the data level, we assume a Poisson distribution for landslide counts, while we compare different "sub-asymptotic" distributions for landslide sizes to flexibly model their upper and lower tails. At the latent level, Poisson intensities and the median of the size distribution vary spatially in terms of fixed and random effects, with shared spatial components capturing cross-correlation between landslide counts and sizes. We robustly model spatial dependence using intrinsic conditional autoregressive priors. Our novel models are fitted efficiently using a customized adaptive Markov chain Monte Carlo algorithm. We show that, for our dataset, sub-asymptotic mark distributions provide improved predictions of large landslide sizes compared to more traditional choices. To showcase the benefits of joint occurrence-size models and illustrate their usefulness for risk assessment, we map landslide hazard along major roads.

stat.ME

A modeler's guide to extreme value software

This review paper surveys recent development in software implementations for extreme value analyses since the publication of Stephenson and Gilleland (2006) and Gilleland et al. (2013), here with a focus on numerical challenges. We provide a comparative review by topic and highlight differences in existing routines, along with listing areas where software development is lacking. The online supplement contains two vignettes providing a comparison of implementations of frequentist and Bayesian estimation of univariate extreme value models.

stat.CO

A flexible Bayesian hierarchical modeling framework for spatially dependent peaks-over-threshold data

In this work, we develop a constructive modeling framework for extreme threshold exceedances in repeated observations of spatial fields, based on general product mixtures of random fields possessing light or heavy-tailed margins and various spatial dependence characteristics, which are suitably designed to provide high flexibility in the tail and at sub-asymptotic levels. Our proposed model is akin to a recently proposed Gamma-Gamma model using a ratio of processes with Gamma marginal distributions, but it possesses a higher degree of flexibility in its joint tail structure, capturing strong dependence more easily. We focus on constructions with the following three product factors, whose different roles ensure their statistical identifiability: a heavy-tailed spatially-dependent field, a lighter-tailed spatially-constant field, and another lighter-tailed spatially-independent field. Thanks to the model's hierarchical formulation, inference may be conveniently performed based on Markov chain Monte Carlo methods. We leverage the Metropolis adjusted Langevin algorithm (MALA) with random block proposals for latent variables, as well as the stochastic gradient Langevin dynamics (SGLD) algorithm for hyperparameters, in order to fit our proposed model very efficiently in relatively high spatio-temporal dimensions, while simultaneously censoring non-threshold exceedances and performing spatial prediction at multiple sites. The censoring mechanism is applied to the spatially independent component, such that only univariate cumulative distribution functions have to be evaluated. We explore the theoretical properties of the novel model, and illustrate the proposed methodology by simulation and application to daily precipitation data from North-Eastern Spain measured at about 100 stations over the period 2011-2020.

stat.ME

Spatiotemporal wildfire modeling through point processes with moderate and extreme marks

Accurate spatiotemporal modeling of conditions leading to moderate and large wildfires provides better understanding of mechanisms driving fire-prone ecosystems and improves risk management. We here develop a joint model for the occurrence intensity and the wildfire size distribution by combining extreme-value theory and point processes within a novel Bayesian hierarchical model, and use it to study daily summer wildfire data for the French Mediterranean basin during 1995--2018. The occurrence component models wildfire ignitions as a spatiotemporal log-Gaussian Cox process. Burnt areas are numerical marks attached to points and are considered as extreme if they exceed a high threshold. The size component is a two-component mixture varying in space and time that jointly models moderate and extreme fires. We capture non-linear influence of covariates (Fire Weather Index, forest cover) through component-specific smooth functions, which may vary with season. We propose estimating shared random effects between model components to reveal and interpret common drivers of different aspects of wildfire activity. This leads to increased parsimony and reduced estimation uncertainty with better predictions. Specific stratified subsampling of zero counts is implemented to cope with large observation vectors. We compare and validate models through predictive scores and visual diagnostics. Our methodology provides a holistic approach to explaining and predicting the drivers of wildfire activity and associated uncertainties.

stat.ME

Modeling spatial extremes using normal mean-variance mixtures

Classical models for multivariate or spatial extremes are mainly based upon the asymptotically justified max-stable or generalized Pareto processes. These models are suitable when asymptotic dependence is present, i.e., the joint tail decays at the same rate as the marginal tail. However, recent environmental data applications suggest that asymptotic independence is equally important and, unfortunately, existing spatial models in this setting that are both flexible and can be fitted efficiently are scarce. Here, we propose a new spatial copula model based on the generalized hyperbolic distribution, which is a specific normal mean-variance mixture and is very popular in financial modeling. The tail properties of this distribution have been studied in the literature, but with contradictory results. It turns out that the proofs from the literature contain mistakes. We here give a corrected theoretical description of its tail dependence structure and then exploit the model to analyze a simulated dataset from the inverted Brown-Resnick process, hindcast significant wave height data in the North Sea, and wind gust data in the state of Oklahoma, USA. We demonstrate that our proposed model is flexible enough to capture the dependence structure not only in the tail but also in the bulk.

stat.ME

Exact Simulation of Max-Infinitely Divisible Processes

Max-infinitely divisible (max-id) processes play a central role in extreme-value theory and include the subclass of all max-stable processes. They allow for a constructive representation based on the pointwise maximum of random functions drawn from a Poisson point process defined on a suitable function space. Simulating from a max-id process is often difficult due to its complex stochastic structure, while calculating its joint density in high dimensions is often numerically infeasible. Therefore, exact and efficient simulation techniques for max-id processes are useful tools for studying the characteristics of the process and for drawing statistical inferences. Inspired by the simulation algorithms for max-stable processes, theory and algorithms to generalize simulation approaches tailored for certain flexible (existing or new) classes of max-id processes are presented. Efficient simulation for a large class of models can be achieved by implementing an adaptive rejection sampling scheme to sidestep a numerical integration step in the algorithm. The results of a simulation study highlight that our simulation algorithm works as expected and is highly accurate and efficient, such that it clearly outperforms customary approximate sampling schemes. As a by-product, new max-id models, which can be represented as pointwise maxima of general location-scale mixtures and possess flexible tail dependence structures capturing a wide range of asymptotic dependence scenarios, are also developed.

stat.ME