SearcharxivSearch

arXiv subjects

Xun Huan

Publications and source records attributed to Xun Huan.

At least 19 recordsLinked to original sources

Matching Urban Flood Sensor Placement to Monitoring Objectives Using Bayesian Optimal Experimental Design

Flood-monitoring sensors are often placed according to coverage, access, or expected inundation. However, the value of a measurement depends on the prediction or decision it is intended to inform. Using tRIBS-Urban simulations and a neural-network surrogate of the August 2014 metropolitan Detroit flood, we examine how this learning target changes single-sensor placement. Across 2,576 candidate locations, we compare parameter-oriented optimal experimental design (PO-OED), which values expected information gain (EIG) about model parameters, with goal-oriented optimal experimental design (GO-OED), which values EIG about specified flood predictions. We also examine how parameter EIG evolves during the event, and illustrate that parameter learning translates unevenly into reductions in predictive uncertainty across locations and lead times. Under GO-OED, point-depth targets favor nearby locations, whereas regional-average and regional maximum-depth targets can favor nonlocal locations. Weighted multi-point objectives retain similar broad spatial patterns, although their computed max-EIG locations differ. Public geospatial data further provide illustrative feasibility and contextual classifications for deployment screening. These results show how monitoring objectives shape sensor placement in optimal experimental design, and motivates an objective-first workflow that defines the intended prediction and priorities, applies field-verified restrictions, and ranks locations by EIG.

stat.AP

Variational Goal-Oriented Optimal Experimental Design for Mixed-Distribution Quantities of Interest: Application to Ship Roll Safety

Goal-oriented optimal experimental design (GO-OED) selects experiments according to the expected information gain (EIG) about a quantity of interest (QoI) rather than the full parameter vector. This work develops a variational GO-OED formulation for mixed discrete-continuous QoI laws arising in probabilistic mechanics when thresholding or event-based transformations map a positive-probability set of uncertain inputs to a common value while other inputs produce continuously varying responses. The motivating application is ship roll safety assessment in random waves, where the QoI is the temporal exceedance probability above a prescribed roll-angle threshold. This quantity is zero when no exceedance occurs and varies continuously over positive values otherwise. A purely continuous variational approximation does not dominate a posterior QoI law containing an atom, yielding an infinite Kullback-Leibler divergence and a trivial Barber-Agakov lower bound of $-\infty$. Scoring atom samples using continuous density values instead changes the objective and does not produce a valid lower-bound estimator. We introduce a mixed variational approximation that models the conditional atom probability and continuous component separately, with a normalizing flow used for the latter. An analytical example recovers the correct EIG landscape, while the ship roll application provides stable EIG lower-bound estimates and identifies informative wave conditions for temporal-exceedance-probability inference.

stat.ME

Bayesian Variational System Identification with Weak-Form Residual Likelihoods

We consider system identification for discovering parameterized operators in governing partial differential equations (PDEs) from noisy spatiotemporal data. Building on variational system identification (VSI), which identifies PDEs through Galerkin weak-form residuals, we develop a Bayesian VSI (B-VSI) framework for operator selection, parameter estimation, and uncertainty quantification. The central idea is to define the likelihood directly in weak-form residual space by propagating observation uncertainty through the weak-form residual map. The resulting likelihood captures heteroscedastic and correlated residual errors while avoiding repeated forward PDE solves during inference. For efficient computation, we use lagged-covariance updates that yield generalized least-squares estimates and conjugate posterior approximations when applicable, together with gradient-based and particle-based methods for more general priors and posterior structures. Model-form uncertainty is handled through sequential operator elimination guided by a residual-space Bayesian information criterion. We demonstrate the framework on state-linear and nonlinear PDEs, including the Fokker--Planck equation and a two-field Cahn--Hilliard equation. The results show that B-VSI accurately recovers active operators and coefficients from noisy data, improves robustness relative to classical VSI, and provides posterior uncertainty estimates for coefficients and derived physical quantities.

cs.CE

Amortized Variational Inference for Joint Posterior and Predictive Distributions in Bayesian Uncertainty Quantification

Bayesian predictive inference propagates parameter uncertainty to quantities of interest through the posterior-predictive distribution. In practice, this is typically performed using a two-stage procedure: first approximating the posterior distribution of model parameters, and then propagating posterior samples through the predictive model via Monte Carlo simulation. This sequential workflow can be computationally demanding, particularly for high-fidelity models such as those governed by partial differential equations. We propose a variational Bayesian framework that directly targets the posterior-predictive distribution and jointly learns variational approximations of both the posterior and the corresponding predictive distribution. The formulation introduces a variational upper bound on the Kullback--Leibler divergence together with moment-based regularization terms. The variational distributions are trained in an amortized manner, shifting computational effort to an offline stage and enabling efficient online inference. Numerical experiments ranging from analytical benchmarks to a finite-element solid mechanics problem demonstrate that the proposed method achieves more accurate predictive distributions than conventional two-stage variational inference, while substantially reducing the cost of online predictive inference.

stat.ML

Mean--Variance Risk-Aware Bayesian Optimal Experimental Design for Nonlinear Models

We propose a variance-penalized formulation of Bayesian optimal experimental design for nonlinear models that augments the classical expected utility criterion with a penalty on utility variability, yielding a mean--variance objective that promotes robust experimental performance. To evaluate this objective, we develop Monte Carlo estimators for the expected utility, its second moment, and the resulting utility variance using prior sampling, thereby avoiding explicit posterior sampling. We then derive leading-order bias and variance expressions using conditional delta-method arguments. The objective is optimized using Bayesian optimization with common random samples to reduce noise. Numerical examples, including a linear-Gaussian benchmark, a nonlinear test problem, and contaminant source inversion in diffusion fields, demonstrate that the proposed approach identifies designs with substantially reduced variability while maintaining competitive expected utility.

stat.ME

2026 Roadmap on Artificial Intelligence and Machine Learning for Smart Manufacturing

The evolution of artificial intelligence (AI) and machine learning (ML) is reshaping smart manufacturing by providing new capabilities for efficiency, adaptability, and autonomy across industrial value chains. However, the deployment of AI and ML in industrial settings still faces critical challenges, including the complexity of industrial big data, effective data management, integration with heterogeneous sensing and control systems, and the demand for trustworthy, explainable, and reliable operation in high-stakes industrial environments. In this roadmap, we present a comprehensive perspective on the foundations, applications, and emerging directions of AI and ML in smart manufacturing. It is structured in three parts. The first highlights the foundations and trends that frame the evolution of AI in smart manufacturing. The second focuses on key topics where AI is already enabling advances, including industrial big data analytics, advanced sensing and perception, autonomous systems, additive and laser-based manufacturing, digital twins, robotics, supply chain and logistics optimization, and sustainable manufacturing. The third section explores non-traditional ML approaches that are opening new frontiers, such as physics-informed AI, generative AI, semantic AI, advanced digital twins, explainable AI, RAMS, data-centric metrology, LLMs, and foundation models for highly connected and complex manufacturing systems. By identifying both opportunities and remaining barriers across these areas, this roadmap outlines the advances needed in methods, integration strategies, and industrial adoption. We hope this roadmap will serve as a guide for researchers, engineers, and practitioners to accelerate innovation, align academic and industrial priorities, and ensure that AI-driven smart manufacturing delivers reliable, sustainable, and scalable impact for the future of manufacturing ecosystems.

cs.AI

Real-Time Physics-Aware Battery Health Monitoring from Partial Charging Profiles via Physics-Informed Neural Networks

Monitoring battery health is essential for ensuring safe and efficient operation. However, there is an inherent trade-off between assessment speed and diagnostic depth-specifically, between rapid overall health estimation and precise identification of internal degradation states. Capturing detailed internal battery information efficiently remains a major challenge, yet such insights are key to understanding the various degradation mechanisms. To address this, we develop a parameterized physics-informed neural network (P-PINNSPM) over the key aging-related parameter space for a single particle model. The model can accurately predict internal battery variables across the parameter space and identifies internal parameters in about 30 seconds-achieving a 47x speedup over the finite volume method-while maintaining high accuracy. These parameters improve the battery state-of-health (SOH) estimation accuracy by at least 60.61%, compared to models without parameter incorporation. Moreover, they enable extrapolation to unseen SOH levels and support robust estimation across diverse charging profiles and operating conditions. Our results demonstrate the strong potential of physics-informed machine learning to advance real-time, data-efficient, and physics-aware battery management systems.

eess.SY

Bifidelity Karhunen-Lo\`eve Expansion Surrogate with Active Learning for Random Fields

We present a bifidelity Karhunen-Lo\`eve expansion (KLE) surrogate model for field-valued quantities of interest (QoIs) under uncertain inputs. The approach combines the spectral efficiency of the KLE with polynomial chaos expansions (PCEs) to preserve an explicit mapping between input uncertainties and output fields. By coupling inexpensive low-fidelity (LF) simulations that capture dominant response trends with a limited number of high-fidelity (HF) simulations that correct for systematic bias, the proposed method enables accurate and computationally affordable surrogate construction. To further improve surrogate accuracy, we form an active learning strategy that adaptively selects new HF evaluations based on the surrogate's generalization error, estimated via cross-validation and modeled using Gaussian process regression. New HF samples are then acquired by maximizing an expected improvement criterion, targeting regions of high surrogate error. The resulting BF-KLE-AL framework is demonstrated on three examples of increasing complexity: a one-dimensional analytical benchmark, a two-dimensional convection-diffusion system, and a three-dimensional turbulent round jet simulation based on Reynolds-averaged Navier--Stokes (RANS) and enhanced delayed detached-eddy simulations (EDDES). Across these cases, the method achieves consistent improvements in predictive accuracy and sample efficiency relative to single-fidelity and random-sampling approaches.

stat.ML

Optimal Stopping for Sequential Bayesian Experimental Design

Sequential Bayesian experimental design is often formulated as a fixed-horizon policy optimization problem, in which the number of experiments is specified before data collection begins. In practical campaigns, however, additional measurements may provide diminishing information relative to their cost, making termination an integral part of experimental design. Common threshold-based stopping rules are easy to implement but myopic, because they compare the current state with a fixed criterion rather than the expected value of future experiments. This work develops a Bayesian optimal stopping framework for sequential experimental design by treating design and stopping as coupled decisions in a finite-horizon sequential decision problem. We prove that, for any fixed design policy, the optimal stopping rule terminates when the immediate terminal reward is no smaller than the expected continuation value. We then derive a policy-gradient method for learning continuous design policies with value-based stopping. The resulting optimization is challenging because the design policy, continuation value, and stopping boundary are mutually dependent, and na\"ive training can become trapped in early-stopping local optima. To address this difficulty, we introduce a curriculum strategy that gradually transitions from forced continuation to adaptive stopping during training. Numerical studies on a linear-Gaussian benchmark, a nonlinear test case, and a contaminant source detection problem show that the proposed approach learns stable, resource-aware design-stopping policies, with the largest gains in settings with strong sequential dependence.

stat.ME

Bayesian Covariance Uncertainty for Adaptive Pilot-Sampling Termination in Multi-fidelity Uncertainty Quantification

Monte Carlo integration becomes prohibitively expensive when each sample requires a high-fidelity model evaluation. Multi-fidelity uncertainty quantification methods mitigate this by combining estimators from high- and low-fidelity models, preserving unbiasedness while reducing variance under a fixed budget. Constructing such estimators optimally requires the model-output covariance matrix, typically estimated from pilot samples. Too few pilot samples lead to inaccurate covariance estimates and suboptimal estimators, while too many consume budget that could be used for final estimation. We propose a Bayesian framework to quantify covariance uncertainty from pilot samples, incorporating prior knowledge and enabling probabilistic assessments of estimator performance. A central component is a flexible $\gamma$-Gaussian prior that ensures computational tractability and supports efficient posterior projection under additional pilot samples. These tools enable adaptive pilot-sampling termination via an interpretable loss criterion that decomposes variance inefficiency into accuracy and cost components. While demonstrated here in the context of approximate control variates (ACV), the framework generalizes to other multi-fidelity estimators. We validate the approach on a monomial benchmark and a PDE-based Darcy flow problem. Across these tests, our adaptive method demonstrates its value for multi-fidelity estimation under limited pilot budgets and expensive models, achieving variance reduction comparable to baseline estimators with oracle covariance.

stat.ME

Goal-Oriented Sequential Bayesian Experimental Design for Causal Learning

We present GO-CBED, a goal-oriented Bayesian framework for sequential causal experimental design. Unlike conventional approaches that select interventions aimed at inferring the full causal model, GO-CBED directly maximizes the expected information gain (EIG) on user-specified causal quantities of interest, enabling more targeted and efficient experimentation. The framework is both non-myopic, optimizing over entire intervention sequences, and goal-oriented, targeting only model aspects relevant to the causal query. To address the intractability of exact EIG computation, we introduce a variational lower bound estimator, optimized jointly through a transformer-based policy network and normalizing flow-based variational posteriors. The resulting policy enables real-time decision-making via an amortized network. We demonstrate that GO-CBED consistently outperforms existing baselines across various causal reasoning and discovery tasks-including synthetic structural causal models and semi-synthetic gene regulatory networks-particularly in settings with limited experimental budgets and complex causal mechanisms. Our results highlight the benefits of aligning experimental design objectives with specific research goals and of forward-looking sequential planning.

cs.LG

Intelligent data collection for network discrimination in material flow analysis using Bayesian optimal experimental design

Material flow analyses (MFAs) are powerful tools for highlighting resource efficiency opportunities in supply chains. MFAs are often represented as directed graphs, with nodes denoting processes and edges representing mass flows. However, network structure uncertainty -- uncertainty in the presence or absence of flows between nodes -- is common and can compromise flow predictions. While collection of more MFA data can reduce network structure uncertainty, an intelligent data acquisition strategy is crucial to optimize the resources (person-hours and money spent on collecting and purchasing data) invested in constructing an MFA. In this study, we apply Bayesian optimal experimental design (BOED), based on the Kullback-Leibler divergence, to efficiently target high-utility MFA data -- data that minimizes network structure uncertainty. We introduce a new method with reduced bias for estimating expected utility, demonstrating its superior accuracy over traditional approaches. We illustrate these advances with a case study on the U.S. steel sector MFA, where the expected utility of collecting specific single pieces of steel mass flow data aligns with the actual reduction in network structure uncertainty achieved by collecting said data from the United States Geological Survey and the World Steel Association. The results highlight that the optimal MFA data to collect depends on the total amount of data being gathered, making it sensitive to the scale of the data collection effort. Overall, our methods support intelligent data acquisition strategies, accelerating uncertainty reduction in MFAs and enhancing their utility for impact quantification and informed decision-making.

stat.AP

A Multi-fidelity Estimator of the Expected Information Gain for Bayesian Optimal Experimental Design

Optimal experimental design (OED) is a framework that leverages a mathematical model of the experiment to identify optimal conditions for conducting the experiment. Under a Bayesian approach, the design objective function is typically chosen to be the expected information gain (EIG). However, EIG is intractable for nonlinear models and must be estimated numerically. Estimating the EIG generally entails some variant of Monte Carlo sampling, requiring repeated data model and likelihood evaluations $\unicode{x2013}$ each involving solving the governing equations of the experimental physics $\unicode{x2013}$ under different sample realizations. This computation becomes impractical for high-fidelity models. We introduce a novel multi-fidelity EIG (MF-EIG) estimator under the approximate control variate (ACV) framework. This estimator is unbiased with respect to the high-fidelity mean, and minimizes variance under a given computational budget. We achieve this by first reparameterizing the EIG so that its expectations are independent of the data models, a requirement for compatibility with ACV. We then provide specific examples under different data model forms, as well as practical enhancements of sample size optimization and sample reuse techniques. We demonstrate the MF-EIG estimator in two numerical examples: a nonlinear benchmark and a turbulent flow problem involving the calibration of shear-stress transport turbulence closure model parameters within the Reynolds-averaged Navier-Stokes model. We validate the estimator's unbiasedness and observe one- to two-orders-of-magnitude variance reduction compared to existing single-fidelity EIG estimators.

stat.CO

Bayesian Model Selection for Network Discrimination and Risk-informed Decision Making in Material Flow Analysis

Material flow analyses (MFAs) provide insight into supply chain level opportunities for resource efficiency. MFAs can be represented as networks with nodes that represent materials, processes, sectors or locations. MFA network structure uncertainty (i.e., the existence or absence of flows between nodes) is pervasive and can undermine the reliability of the flow predictions. This article investigates MFA network structure uncertainty by proposing candidate node-and-flow structures and using Bayesian model selection to identify the most suitable structures and Bayesian model averaging to quantify the parametric mass flow uncertainty. The results of this holistic approach to MFA uncertainty are used in conjunction with the input-output (I/O) method to make risk-informed resource efficiency recommendation. These techniques are demonstrated using a case study on the U.S. steel sector where 16 candidate structures are considered. Model selection highlights 2 networks as most probable based on data collected from the United States Geological Survey and the World Steel Association. Using the I/O method, we then show that the construction sector accounts for the greatest mean share of domestic U.S. steel industry emissions while the automotive and steel products sectors have the highest mean emissions per unit of steel used in the end-use sectors. The uncertainty in the results is used to analyze which end-use sector should be the focus of demand reduction efforts under different appetites for risk. This article's methods generate holistic and transparent MFA uncertainty that account for structural uncertainty, enabling decisions whose outcomes are more robust to the uncertainty.

stat.AP

A Likelihood-Free Approach to Goal-Oriented Bayesian Optimal Experimental Design

Conventional Bayesian optimal experimental design seeks to maximize the expected information gain (EIG) on model parameters. However, the end goal of the experiment often is not to learn the model parameters, but to predict downstream quantities of interest (QoIs) that depend on the learned parameters. And designs that offer high EIG for parameters may not translate to high EIG for QoIs. Goal-oriented optimal experimental design (GO-OED) thus directly targets to maximize the EIG of QoIs. We introduce LF-GO-OED (likelihood-free goal-oriented optimal experimental design), a computational method for conducting GO-OED with nonlinear observation and prediction models. LF-GO-OED is specifically designed to accommodate implicit models, where the likelihood is intractable. In particular, it builds a density ratio estimator from samples generated from approximate Bayesian computation (ABC), thereby sidestepping the need for likelihood evaluations or density estimations. The overall method is validated on benchmark problems with existing methods, and demonstrated on scientific applications of epidemiology and neural science.

stat.CO

Deep Koopman-based Control of Quality Variation in Multistage Manufacturing Systems

This paper presents a modeling-control synthesis to address the quality control challenges in multistage manufacturing systems (MMSs). A new feedforward control scheme is developed to minimize the quality variations caused by process disturbances in MMSs. Notably, the control framework leverages a stochastic deep Koopman (SDK) model to capture the quality propagation mechanism in the MMSs, highlighted by its ability to transform the nonlinear propagation dynamics into a linear one. Two roll-to-roll case studies are presented to validate the proposed method and demonstrate its effectiveness. The overall method is suitable for nonlinear MMSs and does not require extensive expert knowledge.

eess.SY

Optimal experimental design: Formulations and computations

Questions of `how best to acquire data' are essential to modeling and prediction in the natural and social sciences, engineering applications, and beyond. Optimal experimental design (OED) formalizes these questions and creates computational methods to answer them. This article presents a systematic survey of modern OED, from its foundations in classical design theory to current research involving OED for complex models. We begin by reviewing criteria used to formulate an OED problem and thus to encode the goal of performing an experiment. We emphasize the flexibility of the Bayesian and decision-theoretic approach, which encompasses information-based criteria that are well-suited to nonlinear and non-Gaussian statistical models. We then discuss methods for estimating or bounding the values of these design criteria; this endeavor can be quite challenging due to strong nonlinearities, high parameter dimension, large per-sample costs, or settings where the model is implicit. A complementary set of computational issues involves optimization methods used to find a design; we discuss such methods in the discrete (combinatorial) setting of observation selection and in settings where an exact design can be continuously parameterized. Finally we present emerging methods for sequential OED that build non-myopic design policies, rather than explicit designs; these methods naturally adapt to the outcomes of past experiments in proposing new experiments, while seeking coordination among all experiments to be performed. Throughout, we highlight important open questions and challenges.

stat.ME

Variational Bayesian Optimal Experimental Design with Normalizing Flows

Bayesian optimal experimental design (OED) seeks experiments that maximize the expected information gain (EIG) in model parameters. Directly estimating the EIG using nested Monte Carlo is computationally expensive and requires an explicit likelihood. Variational OED (vOED), in contrast, estimates a lower bound of the EIG without likelihood evaluations by approximating the posterior distributions with variational forms, and then tightens the bound by optimizing its variational parameters. We introduce the use of normalizing flows (NFs) for representing variational distributions in vOED; we call this approach vOED-NFs. Specifically, we adopt NFs with a conditional invertible neural network architecture built from compositions of coupling layers, and enhanced with a summary network for data dimension reduction. We present Monte Carlo estimators to the lower bound along with gradient expressions to enable a gradient-based simultaneous optimization of the variational parameters and the design variables. The vOED-NFs algorithm is then validated in two benchmark problems, and demonstrated on a partial differential equation-governed application of cathodic electrophoretic deposition and an implicit likelihood case with stochastic modeling of aphid population. The findings suggest that a composition of 4--5 coupling layers is able to achieve lower EIG estimation bias, under a fixed budget of forward model runs, compared to previous approaches. The resulting NFs produce approximate posteriors that agree well with the true posteriors, able to capture non-Gaussian and multi-modal features effectively.

cs.LG