SearcharxivSearch

arXiv subjects

Alex Lenkoski

Publications and source records attributed to Alex Lenkoski.

At least 19 recordsLinked to original sources

Simulation and evaluation of local daily temperature and precipitation series derived by stochastic downscaling of ERA5 reanalysis

Reanalysis products such as the ERA5 reanalysis are commonly used as proxies for observed atmospheric conditions. These products are convenient to use due to their global coverage, the large number of available atmospheric variables and the physical consistency between these variables, as well as their relatively high spatial and temporal resolutions. However, despite the continuous improvements in accuracy and increasing spatial and temporal resolutions of reanalysis products, they may not always capture local atmospheric conditions, especially for highly localised variables such as precipitation. This paper proposes a computationally efficient stochastic downscaling of ERA5 temperature and precipitation. The method combines information from ERA5 and surface observations from nearby stations in a non-linear regression framework that combines generalised additive models (GAMs) with regression splines and auto-regressive moving average (ARMA) models to produce realistic time series of local daily temperature and precipitation. Using a wide range of evaluation criteria that address different properties of the data, the proposed framework is shown to improve the representation of local temperature and precipitation compared to ERA5 at over 4000 locations in Europe over a period of more than 70 years.

stat.AP

Combining predictive distributions for time-to-event outcomes in meteorology

Combining forecasts from multiple numerical weather prediction (NWP) models have shown substantial benefit over the use of individual forecast products. Although combination, in a broad sense, is widely used in meteorological forecasting, systematic studies of combination methodology in meteorology are scarce. In this article, we study several combination methods, both state-of-the-art and of our own making, with a particular emphasis on situations where one seeks to predict when a particular event of interest will occur. Such time-to-event forecasts require particular methodology and care. We conduct a careful comparison of the different combination methods through an extensive simulation study, where we investigate the conditions under which the combined forecast will outperform the individual forecasting products. Further, we investigate the performance of the methods in a case-study modelling the time to first hard freeze in Norway and parts of Fennoscandia.

stat.AP

Gaussian copula modeling of extreme cold and weak-wind events over Europe conditioned on winter weather regimes

A transition to renewable energy is needed to mitigate climate change. In Europe, this transition has been led by wind energy, which is one of the fastest growing energy sources. However, energy demand and production are sensitive to meteorological conditions and atmospheric variability at multiple time scales. To accomplish the required balance between these two variables, critical conditions of high demand and low wind energy supply must be considered in the design of energy systems. We describe a methodology for modeling joint distributions of meteorological variables without making any assumptions about their marginal distributions. In this context, Gaussian copulas are used to model the correlated nature of cold and weak-wind events. The marginal distributions are modeled with logistic regressions defining two sets of binary variables as predictors: four large-scale weather regimes and the months of the extended winter season. By applying this framework to ERA5 data, we can compute the joint probabilities of co-occurrence of cold and weak-wind events on a high-resolution grid (0.25 deg). Our results show that a) weather regimes must be considered when modeling cold and weak-wind events, b) it is essential to account for the correlations between these events when modeling their joint distribution, c) we need to analyze each month separately, and d) the highest estimated number of days with compound events are associated with the negative phase of the North Atlantic Oscillation (3 days on average over Finland, Ireland, and Lithuania in January, and France and Luxembourg in February) and the Scandinavian Blocking pattern (3 days on average over Ireland in January and Denmark in February). This information could be relevant for application in sub-seasonal to seasonal forecasts of such events.

physics.ao-ph

Probabilistic prediction of the time to hard freeze using seasonal weather forecasts and survival time methods

Agricultural food production and natural ecological systems depend on a range of seasonal climate indicators that describe seasonal patterns in climatological conditions. This paper proposes a probabilistic forecasting framework for predicting the end of the freeze-free season, or the time to a mean daily near-surface air temperature below 0 $^\circ$C (here referred to as hard freeze). The forecasting framework is based on the multi-model seasonal forecast ensemble provided by the Copernicus Climate Data Store and uses techniques from survival analysis for time-to-event data. The original mean daily temperature forecasts are statistically post-processed with a mean and variance correction of each model system before the time-to-event forecast is constructed. In a case study for a region in Fennoscandia covering Norway for the period 1993-2020, the proposed forecasts are found to outperform a climatology forecast from an observation-based data product at locations where the average predicted time to hard freeze is less than 40 days after the initialization date of the forecast on October 1.

stat.AP

Sovereign Risk Indices and Bayesian Theory Averaging

In economic applications, model averaging has found principal use examining the validity of various theories related to observed heterogeneity in outcomes such as growth, development, and trade.Though often easy to articulate, these theories are imperfectly captured quantitatively. A number of different proxies are often collected for a given theory and the uneven nature of this collection requires care when employing model averaging. Furthermore, if valid, these theories ought to be relevant outside of any single narrowly focused outcome equation. We propose a methodology which treats theories as represented by latent indices, these latent processes controlled by model averaging on the proxy level. To achieve generalizability of the theory index our framework assumes a collection of outcome equations. We accommodate a flexible set of generalized additive models, enabling non-Gaussian outcomes to be included. Furthermore, selection of relevant theories also occurs on the outcome level, allowing for theories to be differentially valid. Our focus is on creating a set of theory-based indices directed at understanding a country's potential risk of macroeconomic collapse. These Sovereign Risk Indices are calibrated across a set of different "collapse" criteria, including default on sovereign debt, heightened potential for high unemployment or inflation and dramatic swings in foreign exchange values. The goal of this exercise is to render a portable set of country/year theory indices which can find more general use in the research community.

stat.AP

Rapid adjustment and post-processing of temperature forecast trajectories

Modern weather forecasts are commonly issued as consistent multi-day forecast trajectories with a time resolution of 1-3 hours. Prior to issuing, statistical post-processing is routinely used to correct systematic errors and misrepresentations of the forecast uncertainty. However, once the forecast has been issued, it is rarely updated before it is replaced in the next forecast cycle of the numerical weather prediction (NWP) model. This paper shows that the error correlation structure within the forecast trajectory can be utilized to substantially improve the forecast between the NWP forecast cycles by applying additional post-processing steps each time new observations become available. The proposed rapid adjustment is applied to temperature forecast trajectories from the UK Met Office's convective-scale ensemble MOGREPS-UK. MOGREPS-UK is run four times daily and produces hourly forecasts for up to 36 hours ahead. Our results indicate that the rapidly adjusted forecast from the previous NWP forecast cycle outperforms the new forecast for the first few hours of the next cycle, or until the new forecast itself can be rapidly adjusted, suggesting a new strategy for updating the forecast cycle.

stat.AP

Multivariate postprocessing methods for high-dimensional seasonal weather forecasts

Seasonal weather forecasts are crucial for long-term planning in many practical situations and skillful forecasts may have substantial economic and humanitarian implications. Current seasonal forecasting models require statistical postprocessing of the output to correct systematic biases and unrealistic uncertainty assessments. We propose a multivariate postprocessing approach utilizing covariance tapering, combined with a dimension reduction step based on principal component analysis for efficient computation. Our proposed technique can correctly and efficiently handle non-stationary, non-isotropic and negatively correlated spatial error patterns, and is applicable on a global scale. Further, a moving average approach to marginal postprocessing is shown to flexibly handle trends in biases caused by global warming, and short training periods. In an application to global sea surface temperature forecasts issued by the Norwegian Climate Prediction Model (NorCPM), our proposed methodology is shown to outperform known reference methods.

stat.ME

Probabilistic Forecasting of Temporal Trajectories of Regional Power Production - Part 1: Wind

Renewable energy sources provide a constantly increasing contribution to the total energy production worldwide. However, the power generation from these sources is highly variable due to their dependence on meteorological conditions. Accurate forecasts for the production at various temporal and spatial scales are thus needed for an efficiently operating electricity market. In this article - part 1 - we propose fully probabilistic prediction models for spatially aggregated wind power production at an hourly time scale with lead times up to several days using weather forecasts from numerical weather prediction systems as covariates. After an appropriate cubic transformation of the power production, we build up a multivariate Gaussian prediction model under a Bayesian inference framework which incorporates the temporal error correlation. In an application to predict wind production in Germany, the method provides calibrated and skillful forecasts. Comparison is made between several formulations of the correlation structure.

stat.AP

Probabilistic Forecasting of Temporal Trajectories of Regional Power Production - Part 2: Photovoltaic Solar

We propose a fully probabilistic prediction model for spatially aggregated solar photovoltaic (PV) power production at an hourly time scale with lead times up to several days using weather forecasts from numerical weather prediction systems as covariates. After an appropriate logarithmic transformation of the power production, we develop a multivariate Gaussian prediction model under a Bayesian inference framework. The model incorporates the temporal error correlation yielding physically consistent forecast trajectories. Several formulations of the correlation structure are proposed and investigated. Our method is one of a few approaches that issue full predictive distributions for PV power production. In a case study of PV power production in Germany, the method gives calibrated and skillful forecasts.

stat.AP

Using published bid/ask curves to error dress spot electricity price forecasts

Accurate forecasts of electricity spot prices are essential to the daily operational and planning decisions made by power producers and distributors. Typically, point forecasts of these quantities suffice, particularly in the Nord Pool market where the large quantity of hydro power leads to price stability. However, when situations become irregular, deviations on the price scale can often be extreme and difficult to pinpoint precisely, which is a result of the highly varying marginal costs of generating facilities at the edges of the load curve. In these situations it is useful to supplant a point forecast of price with a distributional forecast, in particular one whose tails are adaptive to the current production regime. This work outlines a methodology for leveraging published bid/ask information from the Nord Pool market to construct such adaptive predictive distributions. Our methodology is a non-standard application of the concept of error-dressing, which couples a feature driven error distribution in volume space with a non-linear transformation via the published bid/ask curves to obtain highly non-symmetric, adaptive price distributions. Using data from the Nord Pool market, we show that our method outperforms more standard forms of distributional modeling. We further show how such distributions can be used to render `warning systems' that issue reliable probabilities of prices exceeding various important thresholds.

q-fin.ST

Exact formulas for the normalizing constants of Wishart distributions for graphical models

Gaussian graphical models have received considerable attention during the past four decades from the statistical and machine learning communities. In Bayesian treatments of this model, the G-Wishart distribution serves as the conjugate prior for inverse covariance matrices satisfying graphical constraints. While it is straightforward to posit the unnormalized densities, the normalizing constants of these distributions have been known only for graphs that are chordal, or decomposable. Up until now, it was unknown whether the normalizing constant for a general graph could be represented explicitly, and a considerable body of computational literature emerged that attempted to avoid this apparent intractability. We close this question by providing an explicit representation of the G-Wishart normalizing constant for general graphs.

math.ST

Spatially adaptive, Bayesian estimation for probabilistic temperature forecasts

Uncertainty in the prediction of future weather is commonly assessed through the use of forecast ensembles that employ a numerical weather prediction model in distinct variants. Statistical postprocessing can correct for biases in the numerical model and improves calibration. We propose a Bayesian version of the standard ensemble model output statistics (EMOS) postprocessing method, in which spatially varying bias coefficients are interpreted as realizations of Gaussian Markov random fields. Our Markovian EMOS (MEMOS) technique utilizes the recently developed stochastic partial differential equation (SPDE) and integrated nested Laplace approximation (INLA) methods for computationally efficient inference. The MEMOS approach shows good predictive performance in a comparative study of 24-hour ahead temperature forecasts over Germany based on the 50-member ensemble of the European Centre for Medium-Range Weather Forecasting (ECMWF).

stat.AP

Efficient sampling of Gaussian graphical models using conditional Bayes factors

Bayesian estimation of Gaussian graphical models has proven to be challenging because the conjugate prior distribution on the Gaussian precision matrix, the G-Wishart distribution, has a doubly intractable partition function. Recent developments provide a direct way to sample from the G-Wishart distribution, which allows for more efficient algorithms for model selection than previously possible. Still, estimating Gaussian graphical models with more than a handful of variables remains a nearly infeasible task. Here, we propose two novel algorithms that use the direct sampler to more efficiently approximate the posterior distribution of the Gaussian graphical model. The first algorithm uses conditional Bayes factors to compare models in a Metropolis-Hastings framework. The second algorithm is based on a continuous time Markov process. We show that both algorithms are substantially faster than state-of-the-art alternatives. Finally, we show how the algorithms may be used to simultaneously estimate both structural and functional connectivity between subcortical brain regions using resting-state fMRI.

q-bio.NC

Bayesian hierarchical modeling of extreme hourly precipitation in Norway

Spatial maps of extreme precipitation are a critical component of flood estimation in hydrological modeling, as well as in the planning and design of important infrastructure. This is particularly relevant in countries such as Norway that have a high density of hydrological power generating facilities and are exposed to significant risk of infrastructure damage due to flooding. In this work, we estimate a spatially coherent map of the distribution of extreme hourly precipitation in Norway, in terms of return levels, by linking generalized extreme value (GEV) distributions with latent Gaussian fields in a Bayesian hierarchical model. Generalized linear models on the parameters of the GEV distribution are able to incorporate location-specific geographic and meteorological information and thereby accommodate these effects on extreme precipitation. A Gaussian field on the GEV parameters captures additional unexplained spatial heterogeneity and overcomes the sparse grid on which observations are collected. We conduct an extensive analysis of the factors that affect the GEV parameters and show that our combination is able to appropriately characterize both the spatial variability of the distribution of extreme hourly precipitation in Norway, and the associated uncertainty in these estimates.

stat.AP

Shape from Texture using Locally Scaled Point Processes

Shape from texture refers to the extraction of 3D information from 2D images with irregular texture. This paper introduces a statistical framework to learn shape from texture where convex texture elements in a 2D image are represented through a point process. In a first step, the 2D image is preprocessed to generate a probability map corresponding to an estimate of the unnormalized intensity of the latent point process underlying the texture elements. The latent point process is subsequently inferred from the probability map in a non-parametric, model free manner. Finally, the 3D information is extracted from the point pattern by applying a locally scaled point process model where the local scaling function represents the deformation caused by the projection of a 3D surface onto a 2D image.

stat.AP

Bayesian Motion Estimation for Dust Aerosols

Dust storms in the earth's major desert regions significantly influence microphysical weather processes, the CO$_2$-cycle and the global climate in general. Recent increases in the spatio-temporal resolution of remote sensing instruments have created new opportunities to understand these phenomena. However, the scale of the data collected and the inherent stochasticity of the underlying process pose significant challenges, requiring a careful combination of image processing and statistical techniques. In particular, using satellite imagery data, we develop a statistical model of atmospheric transport that relies on a latent Gaussian Markov random field (GMRF) for inference. In doing so, we make a link between the optical flow method of Horn and Schunck and the formulation of the transport process as a latent field in a generalized linear model, which enables the use of the integrated nested Laplace approximation for inference. This framework is specified such that it satisfies the so-called integrated continuity equation, thereby intrinsically expressing the divergence of the field as a multiplicative factor covering air compressibility and satellite column projection. The importance of this step -- as well as treating the problem in a fully statistical manner -- is emphasized by a simulation study where inference based on this latent GMRF clearly reduces errors of the estimated flow field. We conclude with a study of the dynamics of dust storms formed over Saharan Africa and show that our methodology is able to accurately and coherently track the storm movement, a critical problem in this field.

stat.AP

A Direct Sampler for G-Wishart Variates

The G-Wishart distribution is the conjugate prior for precision matrices that encode the conditional independencies of a Gaussian graphical model. While the distribution has received considerable attention, posterior inference has proven computationally challenging, in part due to the lack of a direct sampler. In this note, we rectify this situation. The existence of a direct sampler offers a host of new possibilities for the use of G-Wishart variates. We discuss one such development by outlining a new transdimensional model search algorithm--which we term double reversible jump--that leverages this sampler to avoid normalizing constant calculation when comparing graphical models. We conclude with two short studies meant to investigate our algorithm's validity.

stat.CO

A Multivariate Graphical Stochastic Volatility Model

The Gaussian Graphical Model (GGM) is a popular tool for incorporating sparsity into joint multivariate distributions. The G-Wishart distribution, a conjugate prior for precision matrices satisfying general GGM constraints, has now been in existence for over a decade. However, due to the lack of a direct sampler, its use has been limited in hierarchical Bayesian contexts, relegating mixing over the class of GGMs mostly to situations involving standard Gaussian likelihoods. Recent work, however, has developed methods that couple model and parameter moves, first through reversible jump methods and later by direct evaluation of conditional Bayes factors and subsequent resampling. Further, methods for avoiding prior normalizing constant calculations--a serious bottleneck and source of numerical instability--have been proposed. We review and clarify these developments and then propose a new methodology for GGM comparison that blends many recent themes. Theoretical developments and computational timing experiments reveal an algorithm that has limited computational demands and dramatically improves on computing times of existing methods. We conclude by developing a parsimonious multivariate stochastic volatility model that embeds GGM uncertainty in a larger hierarchical framework. The method is shown to be capable of adapting to the extreme swings in market volatility experienced in 2008 after the collapse of Lehman Brothers, offering considerable improvement in posterior predictive distribution calibration.

stat.CO