SearcharxivSearch

arXiv subjects

Michael L. Stein

Publications and source records attributed to Michael L. Stein.

At least 19 recordsLinked to original sources

Optimization of ReaxFF parameters for the $\mathrm{Mo-S}$ system using random optimization and coordinate search

ReaxFF is a molecular dynamics method that can be considered a good approximation to quantum methods for investigating reactive molecular systems consisting of ten thousand to one hundred thousand atoms. While ReaxFF is usually a much faster alternative to quantum methods, the force field consists of nearly 100 parameters per element, which makes the force field development a high dimensional optimization problem. In addition to the high-dimensionality, non-convexity and non-continuity make it a hard problem to optimize. We use random optimization along with coordinate search strategies to optimize efficiently and sample new parameter points that yield good molecular properties close to predefined `reference values' obtained from quantum mechanical methods for the $\mathrm{Mo-S}$ system. We also provide empirical error guaranties starting from any random sample of inputs. We discover new points for the $\mathrm{Mo-S}$ system at adjusted error levels of $13{,}000$ as compared to Sengul et al. (2022) at $70{,}000$ levels under the same loss function, registering over $80\%$ improvement. We also extend our algorithm to an out-of-sample system, $\mathrm{W-S}$, with no training data to record over $70\%$ improvement over Sengul et al. (2021).

stat.AP

An Interpretable Low-Rank State-Space Model for Multi-Horizon Simulation of Large-Scale Regional Temperature Fields

We propose an interpretable low-rank state-space model for conditional simulation of daily regional temperature fields. The central statistical question is whether empirical orthogonal functions (EOFs) can be treated as stable large-scale features of the field rather than only as sample-dependent basis vectors for dimension reduction. The framework represents the dominant temperature field through a small number of retained EOF coefficients, models their seasonal mean and variance structure, propagates the resulting low-dimensional state with stable multivariate dynamics, and uses structured innovations to represent the remaining uncertainty. The resulting reduced-rank model is interpretable, computationally efficient, and suitable for iterative ensemble generation. It is designed to support both inference on the structure of the retained temperature state and prediction through multi-horizon probabilistic field simulation.

stat.ME

Stochastic weather generators for high-frequency wind vector time series

Surface winds can vary substantially from one minute to the next, so there is scope for studying its variation on this fine time scale. Restricting to the month of June to minimize seasonality, this work develops a range of machine learning models for generating realistic time series of surface wind vectors at a site in Lamont, Oklahoma based on more than 30 years of high quality measurements at the minute time scale. Such a generator could be used as an input into models from a range of disciplines, notably for wind energy, but also wildfire spread and aviation, among others. The data show complex diurnal structures in both wind speed and direction that would be challenging to capture with standard time series models, so we consider a number of machine learning approaches to producing a stochastic wind generator based on time vector-quantized variational autoencoders. We consider generating a day's worth of data at a time and generating a day of wind vectors conditional on the previous day's winds. We also study methods for incorporating a discrete weather state variable in the generator. We evaluate the generators using a wide range of formal and informal methods. The best of these generators can capture many but not all of the complex features present in the observational data. In particular, the best of our approaches accurately mimic diurnal changes in wind volatility but struggle to match the observed distribution of extreme wind speeds.

stat.AP

Consistent Infill Estimability of the Regression Slope Between Gaussian Random Fields Under Spatial Confounding

The problem of estimating the slope parameter in regression between two spatial processes under confounding by an unmeasured spatial process has received widespread attention in the recent statistical literature. Yet, a fundamental question remains unresolved: when is this slope consistently estimable under spatial confounding, with existing insights being largely empirical or estimator-specific. We characterize conditions for consistent estimability of the regression slope between Gaussian random fields (GRFs), the common stochastic model for spatial processes, under spatial confounding. Under fixed-domain (infill) asymptotics, we give sufficient conditions for consistent estimability in terms of the smoothness or local behavior of the exposure and confounder processes. When estimability holds, we provide consistent estimators of the slope using local differencing (taking discrete differences or Laplacians of the processes of suitable order). Using functional analysis results on Paley-Wiener spaces, we then provide an easy-to-verify necessary condition for consistent estimability of the slope in terms of the relative spectral tail decays of the confounder and exposure. As a by-product, we establish a novel and general spectral condition on the equivalence of measures on the paths of multivariate GRFs with component fields of varying smoothnesses. We show that for many covariance classes like the Matérn, power-exponential, generalized Cauchy, and coregionalization families, the necessary and sufficient conditions become identical, thereby providing a sharp characterization of consistent estimability of the slope for these processes. The results are extended to multivariate slopes, to accommodate measurement error, to popular classes of non-stationary Gaussian random fields and some non-Gaussian random fields, and for irregular designs.

math.ST

Fixed and Increasing Domain Asymptotics for the Roughness and Scale of Isotropic Gaussian Random Fields

We establish a rigorous asymptotic theory for the joint estimation of roughness and scale parameters in two-dimensional Gaussian random fields with power-law generalized covariances \cite{Matheron1973, Stein1999, Yaglom1987}. Our main results are bivariate central limit theorems for a class of method-of-moments estimators under increasing-domain and fixed-domain asymptotics. The fixed-domain result follows immediately from the increasing-domain result from the self-similarity of Gaussian random fields with power-law generalized covariances \cite{IstasLang1997, Coeurjolly2001, ZhuStein2002}. These results provide a unified distributional framework across these two classical regimes \cite{AvramLeonenkoSakhno2010-ESAIM, BiermeBonamiLeon2011-EJP} that makes the unusual behavior of the estimates under fixed-domain asymptotics intuitively obvious. Our increasing-domain asymptotic results use spatial averages of quadratic forms of (iterated) bilinear product difference filters that yield explicit expressions for the estimates of roughness and scale to which existing theorems on such averages \cite{BreuerMajor1983,Hannan1970} can be readily applied. We further show that the asymptotics remain valid under modestly irregular sampling due to jitter or missing observations. For the fixed-domain setting, the results extend to models that behave sufficiently like the power-law model at high frequencies such as the often used Matérn model \cite{ZhuStein2006, WangLoh2011EJS, KaufmanShaby2017EJS}.

math.ST

Modeling spatial asymmetries in teleconnected extreme temperatures

Combining strengths from deep learning and extreme value theory can help describe complex relationships between variables where extreme events have significant impacts (e.g., environmental or financial applications). Neural networks learn complicated nonlinear relationships from large datasets under limited parametric assumptions. By definition, the number of occurrences of extreme events is small, which limits the ability of the data-hungry, nonparametric neural network to describe rare events. Inspired by recent extreme cold winter weather events in North America caused by atmospheric blocking, we examine several probabilistic generative models for the entire multivariate probability distribution of daily boreal winter surface air temperature. We propose metrics to measure spatial asymmetries, such as long-range anticorrelated patterns that commonly appear in temperature fields during blocking events. Compared to vine copulas, the statistical standard for multivariate copula modeling, deep learning methods show improved ability to reproduce complicated asymmetries in the spatial distribution of ERA5 temperature reanalysis, including the spatial extent of in-sample extreme events.

stat.AP

A Scalable Method to Exploit Screening in Gaussian Process Models with Noise

A common approach to approximating Gaussian log-likelihoods at scale exploits the fact that precision matrices can be well-approximated by sparse matrices in some circumstances. This strategy is motivated by the \emph{screening effect}, which refers to the phenomenon in which the linear prediction of a process $Z$ at a point $\mathbf{x}_0$ depends primarily on measurements nearest to $\mathbf{x}_0$. But simple perturbations, such as i.i.d. measurement noise, can significantly reduce the degree to which this exploitable phenomenon occurs. While strategies to cope with this issue already exist and are certainly improvements over ignoring the problem, in this work we present a new one based on the EM algorithm that offers several advantages. While in this work we focus on the application to Vecchia's approximation (1988), a particularly popular and powerful framework in which we can demonstrate true second-order optimization of M steps, the method can also be applied using entirely matrix-vector products, making it applicable to a very wide class of precision matrix-based approximation methods.

stat.ME

Teleconnected warm and cold extremes of North American wintertime temperatures

Current models for spatial extremes are concerned with the joint upper (or lower) tail of the distribution at two or more locations. Such models cannot account for teleconnection patterns of two-meter surface air temperature ($T_{2m}$) in North America, where very low temperatures in the contiguous Unites States (CONUS) may coincide with very high temperatures in Alaska in the wintertime. This dependence between warm and cold extremes motivates the need for a model with opposite-tail dependence in spatial extremes. This work develops a statistical modeling framework which has flexible behavior in all four pairings of high and low extremes at pairs of locations. In particular, we use a mixture of rotations of common Archimedean copulas to capture various combinations of four-corner tail dependence. We study teleconnected $T_{2m}$ extremes using ERA5 reanalysis of daily average two-meter temperature during the boreal winter. The estimated mixture model quantifies the strength of opposite-tail dependence between warm temperatures in Alaska and cold temperatures in the midlatitudes of North America, as well as the reverse pattern. These dependence patterns are shown to correspond to blocked and zonal patterns of mid-tropospheric flow. This analysis extends the classical notion of correlation-based teleconnections to considering dependence in higher quantiles.

stat.ME

Scalable Computations for Nonstationary Gaussian Processes

Nonstationary Gaussian process models can capture complex spatially varying dependence structures in spatial datasets. However, the large number of observations in modern datasets makes fitting such models computationally intractable with conventional dense linear algebra. In addition, derivative-free or even first-order optimization methods can be slow to converge when estimating many spatially varying parameters. We present here a computational framework that couples an algebraic block-diagonal plus low-rank covariance matrix approximation with stochastic trace estimation to facilitate the efficient use of second-order solvers for maximum likelihood estimation of Gaussian process models with many parameters. We demonstrate the effectiveness of these methods by simultaneously fitting 192 parameters in the popular nonstationary model of Paciorek and Schervish using 107,600 sea surface temperature anomaly measurements.

stat.CO

Fitting Matérn Smoothness Parameters Using Automatic Differentiation

The Matérn covariance function is ubiquitous in the application of Gaussian processes to spatial statistics and beyond. Perhaps the most important reason for this is that the smoothness parameter $ν$ gives complete control over the mean-square differentiability of the process, which has significant implications for the behavior of estimated quantities such as interpolants and forecasts. Unfortunately, derivatives of the Matérn covariance function with respect to $ν$ require derivatives of the modified second-kind Bessel function $\mathcal{K}_ν$ with respect to $ν$. While closed form expressions of these derivatives do exist, they are prohibitively difficult and expensive to compute. For this reason, many software packages require fixing $ν$ as opposed to estimating it, and all existing software packages that attempt to offer the functionality of estimating $ν$ use finite difference estimates for $\partial_ν\mathcal{K}_ν$. In this work, we introduce a new implementation of $\mathcal{K}_ν$ that has been designed to provide derivatives via automatic differentiation (AD), and whose resulting derivatives are significantly faster and more accurate than those computed using finite differences. We provide comprehensive testing for both speed and accuracy and show that our AD solution can be used to build accurate Hessian matrices for second-order maximum likelihood estimation in settings where Hessians built with finite difference approximations completely fail.

stat.CO

Vecchia Likelihood Approximation for Accurate and Fast Inference in Intractable Spatial Extremes Models

Max-stable processes are the most popular models for high-impact spatial extreme events, as they arise as the only possible limits of spatially-indexed block maxima. However, likelihood inference for such models suffers severely from the curse of dimensionality, since the likelihood function involves a combinatorially exploding number of terms. In this paper, we propose using the Vecchia approximation, which conveniently decomposes the full joint density into a linear number of low-dimensional conditional density terms based on well-chosen conditioning sets designed to improve and accelerate inference in high dimensions. Theoretical asymptotic relative efficiencies in the Gaussian setting and simulation experiments in the max-stable setting show significant efficiency gains and computational savings using the Vecchia likelihood approximation method compared to traditional composite likelihoods. Our application to extreme sea surface temperature data at more than a thousand sites across the entire Red Sea further demonstrates the superiority of the Vecchia likelihood approximation for fitting complex models with intractable likelihoods, delivering significantly better results than traditional composite likelihoods, and accurately capturing the extremal dependence structure at lower computational cost.

stat.ME

Nonstationary seasonal model for daily mean temperature distribution bridging bulk and tails

In traditional extreme value analysis, the bulk of the data is ignored, and only the tails of the distribution are used for inference. Extreme observations are specified as values that exceed a threshold or as maximum values over distinct blocks of time, and subsequent estimation procedures are motivated by asymptotic theory for extremes of random processes. For environmental data, nonstationary behavior in the bulk of the distribution, such as seasonality or climate change, will also be observed in the tails. To accurately model such nonstationarity, it seems natural to use the entire dataset rather than just the most extreme values. It is also common to observe different types of nonstationarity in each tail of a distribution. Most work on extremes only focuses on one tail of a distribution, but for temperature, both tails are of interest. This paper builds on a recently proposed parametric model for the entire probability distribution that has flexible behavior in both tails. We apply an extension of this model to historical records of daily mean temperature at several locations across the United States with different climates and local conditions. We highlight the ability of the method to quantify changes in the bulk and tails across the year over the past decades and under different geographic and climatic conditions. The proposed model shows good performance when compared to several benchmark models that are typically used in extreme value analysis of temperature.

stat.AP

Neural Networks for Parameter Estimation in Intractable Models

We propose to use deep learning to estimate parameters in statistical models when standard likelihood estimation methods are computationally infeasible. We show how to estimate parameters from max-stable processes, where inference is exceptionally challenging even with small datasets but simulation is straightforward. We use data from model simulations as input and train deep neural networks to learn statistical parameters. Our neural-network-based method provides a competitive alternative to current approaches, as demonstrated by considerable accuracy and computational time improvements. It serves as a proof of concept for deep learning in statistical parameter estimation and can be extended to other estimation problems.

stat.ME

Linear-Cost Covariance Functions for Gaussian Random Fields

Gaussian random fields (GRF) are a fundamental stochastic model for spatiotemporal data analysis. An essential ingredient of GRF is the covariance function that characterizes the joint Gaussian distribution of the field. Commonly used covariance functions give rise to fully dense and unstructured covariance matrices, for which required calculations are notoriously expensive to carry out for large data. In this work, we propose a construction of covariance functions that result in matrices with a hierarchical structure. Empowered by matrix algorithms that scale linearly with the matrix dimension, the hierarchical structure is proved to be efficient for a variety of random field computations, including sampling, kriging, and likelihood evaluation. Specifically, with $n$ scattered sites, sampling and likelihood evaluation has an $O(n)$ cost and kriging has an $O(\log n)$ cost after preprocessing, particularly favorable for the kriging of an extremely large number of sites (e.g., predicting on more sites than observed). We demonstrate comprehensive numerical experiments to show the use of the constructed covariance functions and their appealing computation time. Numerical examples on a laptop include simulated data of size up to one million, as well as a climate data product with over two million observations.

stat.ME

Flexible nonstationary spatio-temporal modeling of high-frequency monitoring data

Many physical datasets are generated by collections of instruments that make measurements at regular time intervals. For such regular monitoring data, we extend the framework of half-spectral covariance functions to the case of nonstationarity in space and time and demonstrate that this method provides a natural and tractable way to incorporate complex behaviors into a covariance model. Further, we use this method with fully time-domain computations to obtain bona fide maximum likelihood estimators---as opposed to using Whittle-type likelihood approximations, for example---that can still be computed efficiently. We apply this method to very high-frequency Doppler LIDAR vertical wind velocity measurements, demonstrating that the model can expressively capture the extreme nonstationarity of dynamics above and below the atmospheric boundary layer and, more importantly, the interaction of the process dynamics across it.

stat.ME

Scalable Gaussian Process Computations Using Hierarchical Matrices

We present a kernel-independent method that applies hierarchical matrices to the problem of maximum likelihood estimation for Gaussian processes. The proposed approximation provides natural and scalable stochastic estimators for its gradient and Hessian, as well as the expected Fisher information matrix, that are computable in quasilinear $O(n \log^2 n)$ complexity for a large range of models. To accomplish this, we (i) choose a specific hierarchical approximation for covariance matrices that enables the computation of their exact derivatives and (ii) use a stabilized form of the Hutchinson stochastic trace estimator. Since both the observed and expected information matrices can be computed in quasilinear complexity, covariance matrices for MLEs can also be estimated efficiently. After discussing the associated mathematics, we demonstrate the scalability of the method, discuss details of its implementation, and validate that the resulting MLEs and confidence intervals based on the inverse Fisher information matrix faithfully approach those obtained by the exact likelihood.

stat.CO

Locally stationary spatio-temporal interpolation of Argo profiling float data

Argo floats measure seawater temperature and salinity in the upper 2,000 m of the global ocean. Statistical analysis of the resulting spatio-temporal dataset is challenging due to its nonstationary structure and large size. We propose mapping these data using locally stationary Gaussian process regression where covariance parameter estimation and spatio-temporal prediction are carried out in a moving-window fashion. This yields computationally tractable nonstationary anomaly fields without the need to explicitly model the nonstationary covariance structure. We also investigate Student-$t$ distributed fine-scale variation as a means to account for non-Gaussian heavy tails in ocean temperature data. Cross-validation studies comparing the proposed approach with the existing state-of-the-art demonstrate clear improvements in point predictions and show that accounting for the nonstationarity and non-Gaussianity is crucial for obtaining well-calibrated uncertainties. This approach also provides data-driven local estimates of the spatial and temporal dependence scales for the global ocean which are of scientific interest in their own right.

stat.AP

Estimating trends in the global mean temperature record

Given uncertainties in physical theory and numerical climate simulations, the historical temperature record is often used as a source of empirical information about climate change. Many historical trend analyses appear to deemphasize physical and statistical assumptions: examples include regression models that treat time rather than radiative forcing as the relevant covariate and time series methods that account for internal variability nonparametrically. However, given a limited record and the presence of internal variability, estimating radiatively forced historical temperature trends necessarily requires assumptions. Ostensibly empirical methods can involve an inherent conflict in assumptions: they require data records that are short enough for naive trend models to apply but long enough for internal variability to be accounted for. In the context of global mean temperatures, methods that deemphasize assumptions can therefore produce misleading inferences, because the twentieth century trend is complex and the scale of correlation is long relative to the data length. We illustrate how a simple but physically motivated trend model can provide better-fitting and more broadly applicable trend estimates and can address a wider array of questions. The model allows one to distinguish, within a single framework, between uncertainties in the shorter-term versus longer-term response to radiative forcing, with implications not only on historical trends but also on uncertainties in future projections. We also investigate the consequence on inferred uncertainties of the choice of a statistical description of internal variability. While nonparametric methods may seem to avoid making explicit assumptions, we demonstrate how even misspecified parametric methods, if attuned to important characteristics of internal variability, can result in more accurate statements about trend uncertainty.

stat.AP