SearcharxivSearch

arXiv subjects

Richard L. Smith

Publications and source records attributed to Richard L. Smith.

9 recordsLinked to original sources

Fusing Sparse Observations and Dense Simulations for Spatial Extreme Value Analysis: Application to U.S. Coastal Sea Levels

Estimating spatial extremes from sparse observational networks produces uncertain return level maps, but dense output from physics-based simulation models is often available as a complementary data source. We develop a two-stage frequentist frame-work for fusing observations and simulations. In Stage 1, generalized extreme value (GEV) distributions are fitted independently at each site, with a nonstationary location parameter where appropriate to accommodate observed trends. In Stage 2, the parameter estimates from all sources are modeled jointly as a high-dimensional spatial process through a linear model of coregionalization (LMC). Cross-source correlations, estimated from spatially interspersed networks without co-located sites, provide the mechanism for information transfer; an analytic gradient for the resulting likelihood keeps estimation computationally practical. We apply the framework to U.S. coastal sea levels over 1979-2021, fusing 29 NOAA tide gauge records with 100 ADCIRC hydrodynamic simulation sites. Leave-one-out cross-validation shows a 35% reduction in 100-year return level RMSE relative to a gauge-only model. Geographic block cross-validation confirms that fusion benefits persist under spatial extrapolation. The approach is implemented in the R package evfuse.

stat.ME

A Case Study on Quantifying Reliability under Extreme Risk Constraints in Space Missions

In this paper, we employ a Bayesian approach to uncertainty quantification of computer simulations used to assess the probability of rare events. As a case study, we assess the reliability of an Earth reentry capsule for sample return missions that must be able to withstand the reentry loads in order to land intact. Our study uses Gaussian Process modeling under a Bayesian regime to analyze the reentry vehicle's resilience against operational stress. This Bayesian framework allows for a detailed probabilistic evaluation of the system's reliability, indicating our ability to verify stringent safety goals of rare events with a 0.999999 of probability of success. The findings underscore the effectiveness of Bayesian methods for complex uncertainty quantification analyses of computer simulations, providing valuable insights for computational reliability analysis in a risk-averse setting.

stat.AP

ICBM community cancer registry analysis: a focus on Non-Hodgkin Lymphoma cases in missileers

This study investigates the incidence and age at diagnosis of Non-Hodgkin Lymphoma (NHL) among missileers stationed at Malmstrom Air Force Base (MAFB) compared to national benchmarks. The analysis was motivated by reports of elevated cancer diagnoses within the Intercontinental Ballistic Missile (ICBM) community, specifically targeting NHL cases due to initial media focus and data collection through the Torchlight Initiative. The methodology integrates simulation-based estimation of expected diagnoses using incomplete data and expert knowledge on the underlying population. Statistical tests, including the Standardized Incidence Ratio (SIR) and a nonparametric Sign Test, were used to evaluate both the rate of diagnosis and age at diagnosis. The results demonstrate a statistically significant increase in NHL diagnoses among missileers in the later decades, with observed rates surpassing expected benchmarks. The study also finds that the median age of diagnosis is significantly younger for the study population compared to national averages. Key methodological contributions include estimating the population size and service start ages when comprehensive cohort data is unavailable, incorporating uncertainty quantification, and applying multiple hypothesis testing to identify temporal patterns. While the study acknowledges limitations such as small sample size and estimation uncertainty, as well as recognizing the study subjects as self-reported diagnoses, the findings highlight the effectiveness of these statistical techniques in identifying significant deviations from expected rates. Future studies should refine and build upon these methods, incorporating survival analysis and additional covariates to improve the robustness and scope of the results.

q-bio.QM

Alane adsorption and dissociation on the Si(001) surface

We used DFT to study the energetics of the decomposition of alane, AlH3, on the Si(001) surface, as the acceptor complement to PH3. Alane forms a dative bond with the raised atoms of silicon surface dimers, via the Si atom lone pair. We calculated the energies of various structures along the pathway of successive dehydrogenation events following adsorption: AlH2, AlH and Al, finding a gradual, significant decrease in energy. For each stage, we analyse the structure and bonding, and present simulated STM images of the lowest energy structures. Finally, we find that the energy of Al atoms incorporated into the surface, ejecting a Si atom, is comparable to Al adatoms. These findings show that Al incorporation is likely to be as precisely controlled as P incorporation, if slightly less easy to achieve.

cond-mat.mtrl-sci

Air quality and acute deaths in California, 2000-2012

Many studies have sought to determine if there is an association between air quality and acute deaths. Many consider it plausible that current levels of air quality cause acute deaths. However, several factors call causation and even association into question. Observational data sets are large and complex. Multiple testing and multiple modeling can lead to false positive findings. Publication, confirmation and other biases are also possible problems. Moreover, the fact that most data sets used in studies evaluating the relationships among air quality and public health outcomes are not publicly available makes reproducing the claims nearly impossible. Here we have built and made publicly available a dataset containing daily air quality levels, PM2.5 and ozone, daily temperature levels, minimum and maximum and daily relative humidity levels for the eight most populous California air basins. We analyzed the dataset using a moving median analysis, a standard time series analysis, and a prediction analysis within the following analysis strategy. We examine the eight air basins separately to see if estimates replicate across locations. We use leave one year out cross validation analysis to evaluate predictions. Both the moving medians analysis and the standard time series analysis found little evidence for association between air quality and acute deaths. The prediction analysis process was a run as a large factorial design using different models and holding out one year at a time. Among the variables used to predict acute death, most of the daily death variability was explained by time of year or weather variables. In summary, the empirical evidence is that current levels of air quality, ozone and PM2.5, are not causally related to acute deaths for California. An empirical and logical case can be made air quality is not causally related to acute deaths for the rest of the United States.

stat.AP

Dependence Structure of Spatial Extremes Using Threshold Approach

The analysis of spatial extremes requires the joint modeling of a spatial process at a large number of stations and max-stable processes have been developed as a class of stochastic processes suitable for studying spatial extremes. Spatial dependence structure in the extreme value analysis can be measured by max-stable processes. However, there have been few works on the threshold approach of max-stable processes. We propose a threshold version of max-stable process estimation and we apply the pairwise composite likelihood method by Padoan et al. (2010) to estimate spatial dependence parameters. It is of interest to establish limit behavior of the estimates based on the settings of increasing domain asymptotics with stochastic sampling design. Two different types of asymptotic normality are drawn under the second-order regular variation condition for the distribution satisfying the domain of attraction. The theoretical property of dependence parameter estimators in limiting sense is implemented by simulation and a choice of optimal threshold is discussed in this paper.

stat.ME

Approximate Bayesian Computing for Spatial Extremes

Statistical analysis of max-stable processes used to model spatial extremes has been limited by the difficulty in calculating the joint likelihood function. This precludes all standard likelihood-based approaches, including Bayesian approaches. In this paper we present a Bayesian approach through the use of approximate Bayesian computing. This circumvents the need for a joint likelihood function by instead relying on simulations from the (unavailable) likelihood. This method is compared with an alternative approach based on the composite likelihood. We demonstrate that approximate Bayesian computing can result in a lower mean square error than the composite likelihood approach when estimating the spatial dependence of extremes, though at an appreciably higher computational cost. We also illustrate the performance of the method with an application to US temperature data to estimate the risk of crop loss due to an unlikely freeze event.

stat.CO

Pricing Weather Derivatives for Extreme Events

We consider pricing weather derivatives for use as protection against weather extremes. The method described utilizes results from spatial statistics and extreme value theory to first model extremes in the weather as a max-stable process, and then use these models to simulate payments for a general collection of weather derivatives. These simulations capture the spatial dependence of payments. Incorporating results from catastrophe ratemaking, we show how this method can be used to compute risk loads and premiums for weather derivatives which are renewal-additive.

stat.AP

Downscaling extremes: A comparison of extreme value distributions in point-source and gridded precipitation data

There is substantial empirical and climatological evidence that precipitation extremes have become more extreme during the twentieth century, and that this trend is likely to continue as global warming becomes more intense. However, understanding these issues is limited by a fundamental issue of spatial scaling: most evidence of past trends comes from rain gauge data, whereas trends into the future are produced by climate models, which rely on gridded aggregates. To study this further, we fit the Generalized Extreme Value (GEV) distribution to the right tail of the distribution of both rain gauge and gridded events. The results of this modeling exercise confirm that return values computed from rain gauge data are typically higher than those computed from gridded data; however, the size of the difference is somewhat surprising, with the rain gauge data exhibiting return values sometimes two or three times that of the gridded data. The main contribution of this paper is the development of a family of regression relationships between the two sets of return values that also take spatial variations into account. Based on these results, we now believe it is possible to project future changes in precipitation extremes at the point-location level based on results from climate models.

stat.AP