SearcharxivSearch

arXiv subjects

Daniela Castro-Camilo

Publications and source records attributed to Daniela Castro-Camilo.

15 recordsLinked to original sources

Causal Discovery in Multivariate Extremes via Tail Asymmetry

Causal discovery in multivariate extremes is challenging because extreme observations are sparse, dependent, and often affected by latent common shocks. Existing approaches focus on undirected extremal dependence, require prior graph restriction, and do not scale beyond small systems. We introduce tail-induced asymmetry as a principle for causal directionality in heavy-tailed systems, where extreme events propagate asymmetrically so that forward tail prediction is systematically easier than backward prediction. We show that this asymmetry yields identifiable causal direction under a canonical max-linear model and provides a basis for score-based structure learning in the tail regime. Building on this, we propose Sparse Structure diScovery in Multivariate Extremes (S3ME), a two-stage data-driven framework for causal discovery. The first stage performs proxy-adjusted penalized neighbourhood selection to recover a sparse candidate skeleton under latent confounding. The second stage orients edges by minimizing tail prediction risk based on max-linear envelope models, exploiting directional asymmetry. We establish high-dimensional guarantees for skeleton screening and consistency of the score-based estimator under population separation conditions. Simulations demonstrate robustness to latent confounding and favourable scaling relative to existing extremal methods. Applications to river network data and financial tail-risk networks show that the approach recovers sparse, interpretable propagation structures without prespecified graph structure.

stat.ME

Tail-Calibrated Estimation of Extreme Quantile Treatment Effects

Extreme quantile treatment effects (eQTEs) measure the causal impact of a treatment on the tails of an outcome distribution and are central for studying rare, high-impact events. Standard QTE methods often fail in extreme regimes due to data sparsity, while existing eQTE methods rely on restrictive tail assumptions or on interior-quantile theory. We propose the Tail-Calibrated Inverse Estimating Equation (TIEE) framework, which combines information across quantile levels and anchors the tail using extreme value models within a unified estimating equation approach. We establish asymptotic properties of the resulting estimator and evaluate its performance through simulation under different tail behaviours and model misspecifications. An application to extreme precipitation in the Austrian Alps illustrates how TIEE enables observational causal attribution for very rare events under anthropogenic warming. More broadly, the proposed framework establishes a new foundation for causal inference on rare, high-impact outcomes, with relevance across environmental risk, economics, and public health.

stat.ME

Probabilistic forecasting of weather-driven faults in electricity networks: a flexible approach for extreme and non-extreme events

Electricity networks are vulnerable to weather damage, with severe events often leading to faults and power outages. Timely forecasts of fault occurrences, ranging from nowcasts to several days ahead, can enhance preparedness, support faster response, and reduce outage durations. To be operationally useful, such forecasts must quantify uncertainty, enabling risk-informed resource allocation. We present a novel probabilistic framework for forecasting fault counts that captures typical and extreme events. Non-extreme faults are modeled linearly interpolating estimates from multiple additive quantile regressions, while extreme events are described through a discrete generalized Pareto distribution. To incorporate the impact of weather fluctuations, we use ensemble numerical weather predictions, which helps to quantify uncertainty in the forecasts. This approach is designed to provide reliable fault predictions up to four days ahead. We evaluate the model through numerical experiments and apply it to historical fault data from two electricity distribution networks in Great Britain. The resulting forecasts demonstrate substantial improvements over business-as-usual and alternative modeling approaches. A practitioner trial conducted with Scottish Power Energy Networks from October 2024 to March 2025 further demonstrates the operational value of the forecasts. Engineers found them sufficiently reliable to inform decision-making, offering benefits to both network operators and electricity consumers.

stat.AP

XGBoost meets INLA: a two-stage spatio-temporal forecasting of wildfires in Portugal

Wildfires pose a major threat to Portugal, with over 115,000 hectares burned annually on average during 1980-2024, and the country has faced devastating mega-fires such as those in 2017. Accurate forecasts of wildfire occurrence and burned area are therefore essential for firefighting resource allocation and emergency preparedness. In this study, we propose a novel two-stage ensemble that extends the widely used latent Gaussian modelling framework with integrated nested Laplace approximation (INLA) for spatio-temporal wildfire forecasting. Stage 1 applies a gradient boosting model (XGBoost) to environmental covariates and historical fire records to produce one-month-ahead point forecasts of fire counts and burned area. Stage 2 uses these predictions as external covariates in a latent Gaussian model with additional spatiotemporal random effects to generate probabilistic forecasts of monthly total fire counts and burned area at the council level. To capture both moderate and extreme events, we implement the extended generalised Pareto (eGP) likelihood (a sub-asymptotic distribution) within INLA, develop Penalised Complexity (PC) priors for its parameters, and compare the eGP likelihood with common alternatives (e.g., Gamma and Weibull). Our framework tackles the unavailability of future environmental covariates at prediction time and performs strongly for one-month-ahead forecasts.

stat.AP

On the importance of tail assumptions in climate extreme event attribution

Extreme weather events are becoming more frequent and intense, posing serious threats to human life, biodiversity, and ecosystems. A key objective of extreme event attribution (EEA) is to assess whether and to what extent anthropogenic climate change influences such events. Central to EEA is the accurate statistical characterization of atmospheric extremes, which are inherently multivariate or spatial due to their measurement over high-dimensional grids. Within the counterfactual causal inference framework of Pearl, we evaluate how tail assumptions affect attribution conclusions by comparing three multivariate modeling approaches for estimating causation metrics. These include: (i) the multivariate generalized Pareto distribution, which imposes an invariant tail dependence structure; (ii) the factor copula model of Castro-Camilo and Huser (2020), which offers flexible subasymptotic behavior; and (iii) the model of Huser and Wadsworth (2019), which smoothly transitions between different forms of extremal dependence. We assess the implications of these modeling choices in both simulated scenarios (under varying forms of model misspecification) and real data applications, using weekly winter maxima over Europe from the M\'et\'eo-France CNRM model and daily precipitation from the ACCESS-CM2 model over the U.S. Our findings highlight that tail assumptions critically shape causality metrics in EEA. Misspecification of the extremal dependence structure can lead to substantially different and potentially misleading attribution conclusions, underscoring the need for careful model selection and evaluation when quantifying the influence of climate change on extreme events.

stat.AP

Adaptive Bayesian Very Short-Term Wind Power Forecasting Based on the Generalised Logit Transformation

Wind power plays an increasingly significant role in achieving the 2050 Net Zero Strategy. Despite its rapid growth, its inherent variability presents challenges in forecasting. Accurately forecasting wind power generation is one key demand for the stable and controllable integration of renewable energy into existing grid operations. This paper proposes an adaptive method for very short-term forecasting that combines the generalised logit transformation with a Bayesian approach. The generalised logit transformation processes double-bounded wind power data to an unbounded domain, facilitating the application of Bayesian methods. A novel adaptive mechanism for updating the transformation shape parameter is introduced to leverage Bayesian updates by recovering a small sample of representative data. Four adaptive forecasting methods are investigated, evaluating their advantages and limitations through an extensive case study of over 100 wind farms ranging four years in the UK. The methods are evaluated using the Continuous Ranked Probability Score and we propose the use of functional reliability diagrams to assess calibration. Results indicate that the proposed Bayesian method with adaptive shape parameter updating outperforms benchmarks, yielding consistent improvements in CRPS and forecast reliability. The method effectively addresses uncertainty, ensuring robust and accurate probabilistic forecasting which is essential for grid integration and decision-making.

stat.AP

Spatio-temporal fusion of reanalysis and in situ data for censored threshold exceedances of PM2.5

Data fusion models are widely used in air quality monitoring to integrate in situ and large-scale gridded products, offering spatially complete and temporally detailed estimates. However, traditional Gaussian-based models often underestimate extreme pollution values, leading to biased risk assessments. To address this, we present a Bayesian hierarchical data fusion framework rooted in extreme value theory, using the Dirac-delta generalised Pareto distribution to jointly account for threshold and non-threshold exceedances while preserving the timing of exceedance and non-exceedance episodes. Our model is used to describe and predict censored threshold exceedances of PM2.5 pollution in the Greater London region by using CAMS atmospheric composition reanalysis, and in situ observation stations from the automatic urban and rural network (AURN) run by the UK government. Key features of our approach include combining data with varying spatio-temporal resolutions and fully accounting for parameter uncertainties. Results show that our model outperforms Gaussian-based alternatives and standalone reanalysis data in predicting threshold exceedances at the majority of observation sites and can even result in improved spatial patterns of PM2.5 pollution than those discernible from the background data. Moreover, our approach captures greater variability and spatial patterns, such as higher PM2.5 concentrations near coastal areas, which are not evident in the reanalysis data alone.

stat.AP

GPDFlow: Generative Multivariate Threshold Exceedance Modeling via Normalizing Flows

The multivariate generalized Pareto distribution (mGPD) is a common method for modeling extreme threshold exceedance probabilities in environmental and financial risk management. Despite its broad applicability, mGPD faces challenges due to the infinite possible parametrizations of its dependence function, with only a few parametric models available in practice. To address this limitation, we introduce GPDFlow, an innovative mGPD model that leverages normalizing flows to flexibly represent the dependence structure. Unlike traditional parametric mGPD approaches, GPDFlow does not impose explicit parametric assumptions on dependence, resulting in greater flexibility and enhanced performance. Additionally, GPDFlow allows direct inference of marginal parameters, providing insights into marginal tail behavior. We derive tail dependence coefficients for GPDFlow, including a bivariate formulation, a $d$-dimensional extension, and an alternative measure for partial exceedance dependence. A general relationship between the bivariate tail dependence coefficient and the generative samples from normalizing flows is discussed. Through simulations and a practical application analyzing the risk among five major US banks, we demonstrate that GPDFlow significantly improves modeling accuracy and flexibility compared to traditional parametric methods.

stat.ME

A bivariate spatial extreme mixture model for unreplicated heavy metal soil contamination

Geostatistical models for multivariate applications such as heavy metal soil contamination work under Gaussian assumptions and may result in underestimated extreme values and misleading risk assessments (Marchant et al, 2011). A more suitable framework to analyse extreme values is extreme value theory (EVT). However, EVT relies on replications in time, which are generally not available in geochemical datasets. Therefore, using EVT to map soil contamination requires adaptation to be used in the usual single-replicate data framework of soil surveys. We propose a bivariate spatial extreme mixture model to model the body and tail of contaminant pairs, where the tails are described using a stationary generalised Pareto distribution. We demonstrate the performance of our model using a simulation study and through modelling bivariate soil contamination in the Glasgow conurbation. Model results are given as maps of predicted marginal concentrations and probabilities of joint exceedance of soil guideline values. Marginal concentration maps show areas of elevated lead levels along the Clyde River and elevated levels of chromium around the south and southeast villages such as East Kilbride and Wishaw. The joint probability maps show higher probabilities of joint exceedance to the south and southeast of the city centre, following known legacy contamination regions in the Clyde River basin.

stat.AP

A Bayesian multivariate extreme value mixture model

Impact assessment of natural hazards requires the consideration of both extreme and non-extreme events. Extensive research has been conducted on the joint modeling of bulk and tail in univariate settings; however, the corresponding body of research in the context of multivariate analysis is comparatively scant. This study extends the univariate joint modeling of bulk and tail to the multivariate framework. Specifically, it pertains to cases where multivariate observations exceed a high threshold in at least one component. We propose a multivariate mixture model that assumes a parametric model to capture the bulk of the distribution, which is in the max-domain of attraction (MDA) of a multivariate extreme value distribution (mGEVD). The tail is described by the multivariate generalized Pareto distribution, which is asymptotically justified to model multivariate threshold exceedances. We show that if all components exceed the threshold, our mixture model is in the MDA of an mGEVD. Bayesian inference based on multivariate random-walk Metropolis-Hastings and the automated factor slice sampler allows us to incorporate uncertainty from the threshold selection easily. Due to computational limitations, simulations and data applications are provided for dimension $d=2$, but a discussion is provided with views toward scalability based on pairwise likelihood.

stat.ME

Modelling sub-daily precipitation extremes with the blended generalised extreme value distribution

A new method is proposed for modelling the yearly maxima of sub-daily precipitation, with the aim of producing spatial maps of return level estimates. Yearly precipitation maxima are modelled using a Bayesian hierarchical model with a latent Gaussian field, with the blended generalised extreme value (bGEV) distribution used as a substitute for the more standard generalised extreme value (GEV) distribution. Inference is made less wasteful with a novel two-step procedure that performs separate modelling of the scale parameter of the bGEV distribution using peaks over threshold data. Fast inference is performed using integrated nested Laplace approximations (INLA) together with the stochastic partial differential equation (SPDE) approach, both implemented in R-INLA. Heuristics for improving the numerical stability of R-INLA with the GEV and bGEV distributions are also presented. The model is fitted to yearly maxima of sub-daily precipitation from the south of Norway, and is able to quickly produce high-resolution return level maps with uncertainty. The proposed two-step procedure provides an improved model fit over standard inference techniques when modelling the yearly maxima of sub-daily precipitation with the bGEV distribution.

stat.AP

Practical strategies for GEV-based regression models for extremes

The generalised extreme value (GEV) distribution is a three parameter family that describes the asymptotic behaviour of properly renormalised maxima of a sequence of independent and identically distributed random variables. If the shape parameter $ξ$ is zero, the GEV distribution has unbounded support, whereas if $ξ$ is positive, the limiting distribution is heavy-tailed with infinite upper endpoint but finite lower endpoint. In practical applications, we assume that the GEV family is a reasonable approximation for the distribution of maxima over blocks, and we fit it accordingly. This implies that GEV properties, such as finite lower endpoint in the case $ξ>0$, are inherited by the finite-sample maxima, which might not have bounded support. This is particularly problematic when predicting extreme observations based on multiple and interacting covariates. To tackle this usually overlooked issue, we propose a blended GEV distribution, which smoothly combines the left tail of a Gumbel distribution (GEV with $ξ=0$) with the right tail of a Fréchet distribution (GEV with $ξ>0$) and, therefore, has unbounded support. Using a Bayesian framework, we reparametrise the GEV distribution to offer a more natural interpretation of the (possibly covariate-dependent) model parameters. Independent priors over the new location and spread parameters induce a joint prior distribution for the original location and scale parameters. We introduce the concept of property-preserving penalised complexity (P$^3$C) priors and apply it to the shape parameter to preserve first and second moments. We illustrate our methods with an application to NO$_2$ pollution levels in California, which reveals the robustness of the bGEV distribution, as well as the suitability of the new parametrisation and the P$^3$C prior framework.

stat.AP

Bayesian space-time gap filling for inference on extreme hot-spots: an application to Red Sea surface temperatures

We develop a method for probabilistic prediction of extreme value hot-spots in a spatio-temporal framework, tailored to big datasets containing important gaps. In this setting, direct calculation of summaries from data, such as the minimum over a space-time domain, is not possible. To obtain predictive distributions for such cluster summaries, we propose a two-step approach. We first model marginal distributions with a focus on accurate modeling of the right tail and then, after transforming the data to a standard Gaussian scale, we estimate a Gaussian space-time dependence model defined locally in the time domain for the space-time subregions where we want to predict. In the first step, we detrend the mean and standard deviation of the data and fit a spatially resolved generalized Pareto distribution to apply a correction of the upper tail. To ensure spatial smoothness of the estimated trends, we either pool data using nearest-neighbor techniques, or apply generalized additive regression modeling. To cope with high space-time resolution of data, the local Gaussian models use a Markov representation of the Matérn correlation function based on the stochastic partial differential equations (SPDE) approach. In the second step, they are fitted in a Bayesian framework through the integrated nested Laplace approximation implemented in R-INLA. Finally, posterior samples are generated to provide statistical inferences through Monte-Carlo estimation. Motivated by the 2019 Extreme Value Analysis data challenge, we illustrate our approach to predict the distribution of local space-time minima in anomalies of Red Sea surface temperatures, using a gridded dataset (11315 days, 16703 pixels) with artificially generated gaps. In particular, we show the improved performance of our two-step approach over a purely Gaussian model without tail transformations.

stat.ME

A spliced Gamma-Generalized Pareto model for short-term extreme wind speed probabilistic forecasting

Renewable sources of energy such as wind power have become a sustainable alternative to fossil fuel-based energy. However, the uncertainty and fluctuation of the wind speed derived from its intermittent nature bring a great threat to the wind power production stability, and to the wind turbines themselves. Lately, much work has been done on developing models to forecast average wind speed values, yet surprisingly little has focused on proposing models to accurately forecast extreme wind speeds, which can damage the turbines. In this work, we develop a flexible spliced Gamma-Generalized Pareto model to forecast extreme and non-extreme wind speeds simultaneously. Our model belongs to the class of latent Gaussian models, for which inference is conveniently performed based on the integrated nested Laplace approximation method. Considering a flexible additive regression structure, we propose two models for the latent linear predictor to capture the spatio-temporal dynamics of wind speeds. Our models are fast to fit and can describe both the bulk and the tail of the wind speed distribution while producing short-term extreme and non-extreme wind speed probabilistic forecasts.

stat.AP

Local likelihood estimation of complex tail dependence structures, applied to U.S. precipitation extremes

To disentangle the complex non-stationary dependence structure of precipitation extremes over the entire contiguous U.S., we propose a flexible local approach based on factor copula models. Our sub-asymptotic spatial modeling framework yields non-trivial tail dependence structures, with a weakening dependence strength as events become more extreme, a feature commonly observed with precipitation data but not accounted for in classical asymptotic extreme-value models. To estimate the local extremal behavior, we fit the proposed model in small regional neighborhoods to high threshold exceedances, under the assumption of local stationarity, which allows us to gain in flexibility. Adopting a local censored likelihood approach, inference is made on a fine spatial grid, and local estimation is performed by taking advantage of distributed computing resources and the embarrassingly parallel nature of this estimation procedure. The local model is efficiently fitted at all grid points, and uncertainty is measured using a block bootstrap procedure. An extensive simulation study shows that our approach can adequately capture complex, non-stationary dependencies, while our study of U.S. winter precipitation data reveals interesting differences in local tail structures over space, which has important implications on regional risk assessment of extreme precipitation events.

stat.AP