SearcharxivSearch

arXiv subjects

Teresa Portone

Publications and source records attributed to Teresa Portone.

8 recordsLinked to original sources

Duality and Error for Predictively Oriented Inference

Predictively oriented (PrO) inference quantifies uncertainty by selecting a distribution over model parameters to optimize a scoring rule applied to the induced predictive distribution, together with a divergence penalty from a reference distribution. By applying the scoring rule after averaging model densities, PrO inference targets predictive performance, accounting for model misspecification. We focus on the logarithmic score with general $ϕ$-divergence regularization. Our contributions are twofold. First, we derive a finite-dimensional dual formulation of PrO inference. For $n$ observations, the dual problem has $n+1$ variables. We establish zero-duality-gap criteria and optimality conditions that relate the primal and dual solutions. When primal and dual solutions exist, these conditions yield a semi-analytical representation of the PrO posterior and certificates for assessing the accuracy of numerical solutions. For Kullback--Leibler regularization, the posterior has an exponential form. Second, we derive a finite-sample excess predictive-risk bound for approximate PrO posteriors that separates sampling fluctuation, approximation under a divergence budget, regularization, and numerical optimization error. The result applies even when the benchmark predictive risk is not attained by any probability distribution over the model parameters having finite divergence from the reference distribution. We use an exactly solvable categorical example to show that predictive-risk convergence can imply convergence to a unique predictive distribution even though the parameter distributions have no weak limit on the original parameter space. The example also shows that different $ϕ$-divergences can require different regularization schedules. We conclude with a misspecified Gaussian location-mixture example that illustrates the dual computation, primal recovery, and numerical accuracy checks.

stat.ME

Scalable extensions to given-data Sobol' index estimators

Given-data methods for variance-based sensitivity analysis have significantly advanced the feasibility of Sobol' index computation for computationally expensive models and models with many inputs. However, the limitations of existing methods still preclude their application to models with an extremely large number of inputs. In this work, we present practical and theoretical extensions to the existing given-data Sobol' index method, which allow variance-based sensitivity analysis to be efficiently performed on large models such as neural networks, which have $>10^4$ inputs. For models of this size, holding all input-output evaluations simultaneously in memory---as required by existing methods---can quickly become impractical. Our extensions include a general definition of the given-data Sobol' index estimator with arbitrary partition, a streaming algorithm to process input-output samples in batches, and an asymptotic analysis of the new estimator that motivates a practical screening heuristic for small indices. We show that the equiprobable partition employed in existing given-data methods can introduce significant bias into Sobol' index estimates even at large sample sizes and provide numerical analyses that demonstrate why this can occur. We also show that the streaming algorithm can achieve comparable accuracy and runtime while substantially reducing memory requirements, enabling sensitivity analysis of models with much larger input dimension. We demonstrate our novel developments on two application problems in neural network modeling.

stat.ML

Inference in the presence of model-form uncertainties: Leveraging a prediction-oriented approach to improve uncertainty characterization

Bayesian inference is a popular approach to calibrating uncertainties, but it can underpredict such uncertainties when model misspecification is present, impacting its reliability to inform decision making. Recently, the statistics and machine learning communities have developed prediction-oriented inference approaches that provide better calibrated uncertainties and adapt to the level of misspecification present. However, these approaches have yet to be demonstrated in the context of complex scientific applications where phenomena of interest are governed by physics-based models. Such settings often involve single realizations of high-dimensional spatio-temporal data and nonlinear, computationally expensive parameter-to-observable maps. This work investigates variational prediction-oriented inference in problems exhibiting these relevant features; namely, we consider a polynomial model and a contaminant transport problem governed by advection-diffusion equations. The prediction-oriented loss is formulated as the log-predictive probability of the calibration data. We study the effects of increasing misspecification and noise, and we assess approximations of the predictive density using Monte Carlo sampling and component-wise kernel density estimation. A novel aspect of this work is applying prediction-oriented inference to the calibration of model-form uncertainty (MFU) representations, which are embedded physics-based modifications to the governing equations that aim to reduce (but rarely eliminate) model misspecification. The computational results demonstrate that prediction-oriented frameworks can provide better uncertainty characterizations in comparison to standard inference while also being amenable to the calibration of MFU representations.

cs.CE

Closure Term Estimation in Spatiotemporal Models of Dynamical Systems

Closure modeling - the statistical modeling of missing dynamics in the natural sciences and engineering - is a growing and active area of research. Existing methods for closure modeling are often computationally prohibitive, lack uncertainty quantification, or require noise-free observations of the temporal derivatives over the system state. We propose a novel, computationally efficient approach for the modeling and estimation of closure terms over the spatiotemporal domain that provides uncertainty quantification and is effective even when the observations of the system state are sparse or contain moderate levels of noise. The efficacy of our approach is demonstrated in both one and two spatial dimensions through numerical experiments using the Fisher-KPP reaction-diffusion equation and the advection-diffusion equation as exemplars.

stat.ME

Quantifying model prediction sensitivity to model-form uncertainty

Model-form uncertainty (MFU) in assumptions made during physics-based model development is widely considered a significant source of uncertainty; however, there are limited approaches that can quantify MFU in predictions extrapolating beyond available data. As a result, it is challenging to know how important MFU is in practice, especially relative to other sources of uncertainty in a model, making it difficult to prioritize resources and efforts to drive down error in model predictions. To address these challenges, we present a novel method to quantify the importance of uncertainties associated with model assumptions. We combine parameterized modifications to assumptions (called MFU representations) with grouped variance-based sensitivity analysis to measure the importance of assumptions. We demonstrate how, in contrast to existing methods addressing MFU, our approach can be applied without access to calibration data. However, if calibration data is available, we demonstrate how it can be used to inform the MFU representation, and how variance-based sensitivity analysis can be meaningfully applied even in the presence of dependence between parameters (a common byproduct of calibration).

cs.CE

Hybrid Physics-Data Enrichments to Represent Uncertainty in Reduced Gas-Surface Chemistry Models for Hypersonic Flight

During hypersonic flight, air reacts with a planetary re-entry vehicle's thermal protection system (TPS), creating reaction products that deplete the TPS. Reliable assessment of TPS performance depends on accurate ablation models. New finite-rate gas-surface chemistry models are advancing state-of-the-art in TPS ablation modeling, but model reductions that omit chemical species and reactions may be necessary in some cases for computational tractability. This work develops hybrid physics-based and data-driven enrichments to improve the predictive capability and quantify uncertainties in such low-fidelity models while maintaining computational tractability. We focus on discrepancies in predicted carbon monoxide production that arise because the low-fidelity model tracks only a subset of reactions. To address this, we embed targeted enrichments into the low-fidelity model to capture the influence of omitted reactions. Numerical results show that the hybrid enrichments significantly improve predictive accuracy while requiring the addition of only three reactions.

cs.CE

Bayesian inference of an uncertain generalized diffusion operator

This paper defines a novel Bayesian inverse problem to infer an infinite-dimensional uncertain operator appearing in a differential equation, whose action on an observable state variable affects its dynamics. Inference is made tractable by parametrizing the operator using its eigendecomposition. The plausibility of operator inference in the sparse data regime is explored in terms of an uncertain, generalized diffusion operator appearing in an evolution equation for a contaminant's transport through a heterogeneous porous medium. Sparse data are augmented with prior information through the imposition of deterministic constraints on the eigendecomposition and the use of qualitative information about the system in the definition of the prior distribution. Limited observations of the state variable's evolution are used as data for inference, and the dependence on the solution of the inverse problem is studied as a function of the frequency of observations, as well as on whether or not the data is collected as a spatial or time series.

stat.AP

A Stochastic Operator Approach to Model Inadequacy with Applications to Contaminant Transport

The mathematical models used to represent physical phenomena are generally known to be imperfect representations of reality. Model inadequacies arise for numerous reasons, such as incomplete knowledge of the phenomena or computational intractability of more accurate models. In such situations it is impractical or impossible to improve the model, but necessity requires its use to make predictions. With this in mind, it is important to represent the uncertainty that a model's inadequacy causes in its predictions, as neglecting to do so can cause overconfidence in its accuracy. A powerful approach to addressing model inadequacy leverages the composite nature of physical models by enriching a flawed embedded closure model with a stochastic error representation. This work outlines steps in the development of a stochastic operator as an inadequacy representation by establishing the framework for inferring an infinite-dimensional operator and by introducing a novel method for interrogating available high-fidelity models to learn about modeling error.

cs.CE