SearcharxivSearch

arXiv subjects

William Kleiber

Publications and source records attributed to William Kleiber.

13 recordsLinked to original sources

Demographics of Mesoscale Eddies in an Eddy-Permitting Ocean Model and Reanalysis

Ocean mesoscale eddies can be thought of as the "weather" of the ocean and strongly influence the ocean's physics, chemistry, and biology; they influence other components of the Earth system via air-sea and sea-ice interactions, and are crucial drivers of marine heat waves. Thus, proper modeling of eddies in both historical and future climates is crucial to accurately capturing the Earth system. Climate projections using global coupled models with eddying ocean components are only recently starting to be more widely used. Despite their critical role in understanding and forecasting climate characteristics, these so-called eddy-permitting models have not been explored to verify that resolved eddies are realistic, and thus any downstream scientific testing of hypotheses in biogeochemistry, ocean physics or other associated Earth systems impacted by eddies hinge on this critical assumption. This paper compares observed eddies with lifetimes longer than 6 weeks present in $1/4^\circ$ satellite altimetry data with observed eddies in $1/4^\circ$ reanalysis data and ocean model output. When compared to eddies observed in satellite altimetry data, eddies in reanalysis data and ocean model output are missing almost 30% of the number of eddy trajectories. In addition to missing eddy trajectories, the characteristics of eddies in reanalysis data and ocean model output differ from eddies observed in satellite altimetry data. At a high level, eddies in reanalysis data and ocean model output tend to live longer, are larger, and are weaker than eddies in observed altimetry data. This paper presents a variety of statistics describing these differences both spatially and in global aggregate.

physics.ao-ph

Fast Likelihood-Free Parameter Estimation for L\'evy Processes

L\'evy processes are widely used in financial modeling due to their ability to capture discontinuities and heavy tails, which are common in high-frequency asset return data. However, parameter estimation remains a challenge when associated likelihoods are unavailable or costly to compute. We propose a fast and accurate method for L\'evy parameter estimation using the neural Bayes estimation (NBE) framework -- a simulation-based, likelihood-free approach that leverages permutation-invariant neural networks to approximate Bayes estimators. We contribute new theoretical results, showing that NBE results in consistent estimators whose risk converges to the Bayes estimator under mild conditions. Moreover, through extensive simulations across several L\'evy models, we show that NBE outperforms traditional methods in both accuracy and runtime, while also enabling two complementary approaches to uncertainty quantification. We illustrate our approach on a challenging high-frequency cryptocurrency return dataset, where the method captures evolving parameter dynamics and delivers reliable and interpretable inference at a fraction of the computational cost of traditional methods. NBE provides a scalable and practical solution for inference in complex financial models, enabling parameter estimation and uncertainty quantification over an entire year of data in just seconds. We additionally investigate nearly a decade of high-frequency Bitcoin returns, requiring less than one minute to estimate parameters under the proposed approach.

stat.ML

Modeling Large Nonstationary Spatial Data with the Full-Scale Basis Graphical Lasso

We propose a new approach for the modeling large datasets of nonstationary spatial processes that combines a latent low rank process and a sparse covariance model. The low rank component coefficients are endowed with a flexible graphical Gaussian Markov random field model. The utilization of a low rank and compactly-supported covariance structure combines the full-scale approximation and the basis graphical lasso; we term this new approach the full-scale basis graphical lasso (FSBGL). Estimation employs a graphical lasso-penalized likelihood, which is optimized using a difference-of-convex scheme. We illustrate the proposed approach on synthetic fields as well as with a challenging high-resolution simulation dataset of the thermosphere. In a comparison against state-of-the-art spatial models, the FSBGL performs better at capturing salient features of the thermospheric temperature fields, even with limited available training data.

stat.ME

Modeling massive highly-multivariate nonstationary spatial data with the basis graphical lasso

We propose a new modeling framework for highly-multivariate spatial processes that synthesizes ideas from recent multiscale and spectral approaches with graphical models. The basis graphical lasso writes a univariate Gaussian process as a linear combination of basis functions weighted with entries of a Gaussian graphical vector whose graph is estimated from optimizing an $\ell_1$ penalized likelihood. This paper extends the setting to a multivariate Gaussian process where the basis functions are weighted with Gaussian graphical vectors. We motivate a model where the basis functions represent different levels of resolution and the graphical vectors for each level are assumed to be independent. Using an orthogonal basis grants linear complexity and memory usage in the number of spatial locations, the number of basis functions, and the number of realizations. An additional fusion penalty encourages a parsimonious conditional independence structure in the multilevel graphical model. We illustrate our method on a large climate ensemble from the National Center for Atmospheric Research's Community Atmosphere Model that involves 40 spatial processes.

stat.ME

Stochastic Tropical Cyclone Precipitation Field Generation

Tropical cyclones are important drivers of coastal flooding which have severe negative public safety and economic consequences. Due to the rare occurrence of such events, high spatial and temporal resolution historical storm precipitation data are limited in availability. This paper introduces a statistical tropical cyclone space-time precipitation generator given limited information from storm track datasets. Given a handful of predictor variables that are common in either historical or simulated storm track ensembles such as pressure deficit at the storm's center, radius of maximal winds, storm center and direction, and distance to coast, the proposed stochastic model generates space-time fields of quantitative precipitation over the study domain. Statistically novel aspects include that the model is developed in Lagrangian coordinates with respect to the dynamic storm center that uses ideas from low-rank representations along with circular process models. The model is trained on a set of tropical cyclone data from an advanced weather forecasting model over the Gulf of Mexico and southern United States, and is validated by cross-validation. Results show the model appropriately captures spatial asymmetry of cyclone precipitation patterns, total precipitation as well as the local distribution of precipitation at a set of case study locations along the coast.

stat.AP

Nonrigid registration using Gaussian processes and local likelihood estimation

Surface registration, the task of aligning several multidimensional point sets, is a necessary task in many scientific fields. In this work, a novel statistical approach is developed to solve the problem of nonrigid registration. While the application of an affine transformation results in rigid registration, using a general nonlinear function to achieve nonrigid registration is necessary when the point sets require deformations that change over space. The use of a local likelihood-based approach using windowed Gaussian processes provides a flexible way to accurately estimate the nonrigid deformation. This strategy also makes registration of massive data sets feasible by splitting the data into many subsets. The estimation results yield spatially-varying local rigid registration parameters. Gaussian process surface models are then fit to the parameter fields, allowing prediction of the transformation parameters at unestimated locations, specifically at observation locations in the unregistered data set. Applying these transformations results in a global, nonrigid registration. A penalty on the transformation parameters is included in the likelihood objective function. Combined with smoothing of the local estimates from the surface models, the nonrigid registration model can prevent the problem of overfitting. The efficacy of the nonrigid registration method is tested in two simulation studies, varying the number of windows and number of points, as well as the type of deformation. The nonrigid method is applied to a pair of massive remote sensing elevation data sets exhibiting complex geological terrain, with improved accuracy and uncertainty quantification in a cross validation study versus two rigid registration methods.

stat.ME

Penalized basis models for very large spatial datasets

Many modern spatial models express the stochastic variation component as a basis expansion with random coefficients. Low rank models, approximate spectral decompositions, multiresolution representations, stochastic partial differential equations and empirical orthogonal functions all fall within this basic framework. Given a particular basis, stochastic dependence relies on flexible modeling of the coefficients. Under a Gaussianity assumption, we propose a graphical model family for the stochastic coefficients by parameterizing the precision matrix. Sparsity in the precision matrix is encouraged using a penalized likelihood framework. Computations follow from a majorization-minimization approach, a byproduct of which is a connection to the graphical lasso. The result is a flexible nonstationary spatial model that is adaptable to very large datasets. We apply the model to two large and heterogeneous spatial datasets in statistical climatology and recover physically sensible graphical structures. Moreover, the model performs competitively against the popular LatticeKrig model in predictive cross-validation, but substantially improves the Akaike information criterion score.

stat.ME

Improving particle filter performance by smoothing observations

This article shows that increasing the observation variance at small scales can reduce the ensemble size required to avoid collapse in particle filtering of spatially-extended dynamics and improve the resulting uncertainty quantification at large scales. Particle filter weights depend on how well ensemble members agree with observations, and collapse occurs when a few ensemble members receive most of the weight. Collapse causes catastrophic variance underestimation. Increasing small-scale variance in the observation error model reduces the incidence of collapse by de-emphasizing small-scale differences between the ensemble members and the observations. Doing so smooths the posterior mean, though it does not smooth the individual ensemble members. Two options for implementing the proposed observation error model are described. Taking discretized elliptic differential operators as an observation error covariance matrix provides the desired property of a spectrum that grows in the approach to small scales. This choice also introduces structure exploitable by scalable computation techniques, including multigrid solvers and multiresolution approximations to the corresponding integral operator. Alternatively the observations can be smoothed and then assimilated under the assumption of independent errors, which is equivalent to assuming large errors at small scales. The method is demonstrated on a linear stochastic partial differential equation, where it significantly reduces the occurrence of particle filter collapse while maintaining accuracy. It also improves continuous ranked probability scores by as much as 25%, indicating that the weighted ensemble more accurately represents the true distribution. The method is compatible with other techniques for improving the performance of particle filters.

stat.AP

Cross-Covariance Functions for Multivariate Geostatistics

Continuously indexed datasets with multiple variables have become ubiquitous in the geophysical, ecological, environmental and climate sciences, and pose substantial analysis challenges to scientists and statisticians. For many years, scientists developed models that aimed at capturing the spatial behavior for an individual process; only within the last few decades has it become commonplace to model multiple processes jointly. The key difficulty is in specifying the cross-covariance function, that is, the function responsible for the relationship between distinct variables. Indeed, these cross-covariance functions must be chosen to be consistent with marginal covariance functions in such a way that the second-order structure always yields a nonnegative definite covariance matrix. We review the main approaches to building cross-covariance models, including the linear model of coregionalization, convolution methods, the multivariate Matérn and nonstationary and space-time extensions of these among others. We additionally cover specialized constructions, including those designed for asymmetry, compact support and spherical domains, with a review of physics-constrained models. We illustrate select models on a bivariate regional climate model output example for temperature and pressure, along with a bivariate minimum and maximum temperature observational dataset; we compare models by likelihood value as well as via cross-validation co-kriging studies. The article closes with a discussion of unsolved problems.

stat.ME

Coherence for Random Fields

Multivariate spatial field data are increasingly common and whose modeling typically relies on building cross-covariance functions to describe cross-process relationships. An alternative viewpoint is to model the matrix of spectral measures. We develop the notions of coherence, phase and gain for multidimensional stationary processes. Coherence, as a function of frequency, can be seen to be a measure of linear relationship between two spatial processes at that frequency band. We use the coherence function to illustrate fundamental limitations on a number of previously proposed constructions for multivariate processes, suggesting these options are not viable for real data. We also give natural interpretations to cross-covariance parameters of the Matern class, where the smoothness indexes dependence at low frequencies while the range parameter can imply dependence at low or high frequencies. Estimation follows from smoothed multivariate periodogram matrices. We illustrate the estimation and interpretation of these functions on two datasets, forecast and reanalysis sea level pressure and geopotential heights over the equatorial region. Examining these functions lends insight that would otherwise be difficult to detect and model using standard cross-covariance formulations.

math.ST

Parameter tuning for a multi-fidelity dynamical model of the magnetosphere

Geomagnetic storms play a critical role in space weather physics with the potential for far reaching economic impacts including power grid outages, air traffic rerouting, satellite damage and GPS disruption. The LFM-MIX is a state-of-the-art coupled magnetospheric-ionospheric model capable of simulating geomagnetic storms. Imbedded in this model are physical equations for turning the magnetohydrodynamic state parameters into energy and flux of electrons entering the ionosphere, involving a set of input parameters. The exact values of these input parameters in the model are unknown, and we seek to quantify the uncertainty about these parameters when model output is compared to observations. The model is available at different fidelities: a lower fidelity which is faster to run, and a higher fidelity but more computationally intense version. Model output and observational data are large spatiotemporal systems; the traditional design and analysis of computer experiments is unable to cope with such large data sets that involve multiple fidelities of model output. We develop an approach to this inverse problem for large spatiotemporal data sets that incorporates two different versions of the physical model. After an initial design, we propose a sequential design based on expected improvement. For the LFM-MIX, the additional run suggested by expected improvement diminishes posterior uncertainty by ruling out a posterior mode and shrinking the width of the posterior distribution. We also illustrate our approach using the Lorenz `96 system of equations for a simplified atmosphere, using known input parameters. For the Lorenz `96 system, after performing sequential runs based on expected improvement, the posterior mode converges to the true value and the posterior variability is reduced.

stat.AP

Daily minimum and maximum temperature simulation over complex terrain

Spatiotemporal simulation of minimum and maximum temperature is a fundamental requirement for climate impact studies and hydrological or agricultural models. Particularly over regions with variable orography, these simulations are difficult to produce due to terrain driven nonstationarity. We develop a bivariate stochastic model for the spatiotemporal field of minimum and maximum temperature. The proposed framework splits the bivariate field into two components of "local climate" and "weather." The local climate component is a linear model with spatially varying process coefficients capturing the annual cycle and yielding local climate estimates at all locations, not only those within the observation network. The weather component spatially correlates the bivariate simulations, whose matrix-valued covariance function we estimate using a nonparametric kernel smoother that retains nonnegative definiteness and allows for substantial nonstationarity across the simulation domain. The statistical model is augmented with a spatially varying nugget effect to allow for locally varying small scale variability. Our model is applied to a daily temperature data set covering the complex terrain of Colorado, USA, and successfully accommodates substantial temporally varying nonstationarity in both the direct-covariance and cross-covariance functions.

stat.AP