SearcharxivSearch

arXiv subjects

Shixuan Zhang

Publications and source records attributed to Shixuan Zhang.

15 recordsLinked to original sources

Retrofitting Earth System Models with Cadence-Limited Neural Operator Updates

Coarse resolution, imperfect parameterizations, and uncertain initial states and forcings limit Earth-system model (ESM) predictions. Traditional bias correction via data assimilation improves constrained simulations but offers limited benefit once models run freely. We introduce an operator-learning framework that maps instantaneous model states to bias-correction tendencies and applies them online during integration. Building on a U-Net backbone, we develop two operator architectures Inception U-Net (IUNet) and a multi-scale network (M\&M) that combine diverse upsampling and receptive fields to capture multiscale nonlinear features under Energy Exascale Earth System Model (E3SM) runtime constraints. Trained on two years E3SM simulations nudged toward ERA5 reanalysis, the operators generalize across height levels and seasons. Both architectures outperform standard U-Net baselines in offline tests, indicating that functional richness rather than parameter count drives performance. In online hybrid E3SM runs, M\&M delivers the most consistent bias reductions across variables and vertical levels. The ML-augmented configurations remain stable and computationally feasible in multi-year simulations, providing a practical pathway for scalable hybrid modeling. Our framework emphasizes long-term stability, portability, and cadence-limited updates, demonstrating the utility of expressive ML operators for learning structured, cross-scale relationships and retrofitting legacy ESMs.

cs.LG

On Distributionally Robust Multistage Convex Optimization: Data-driven Models and Performance

This paper presents a novel algorithmic study with extensive numerical experiments of distributionally robust multistage convex optimization (DR-MCO). Following the previous work on dual dynamic programming (DDP) algorithmic framework for DR-MCO, we focus on data-driven DR-MCO models with Wasserstein ambiguity sets that allow probability measures with infinite supports. These data-driven Wasserstein DR-MCO models have out-of-sample performance guarantees and adjustable in-sample conservatism. Then by exploiting additional concavity or convexity in the uncertain cost functions, we design exact single stage subproblem oracle (SSSO) implementations that ensure the convergence of DDP algorithms. We test the data-driven Wasserstein DR-MCO models against multistage robust convex optimization (MRCO), risk-neutral and risk-averse multistage stochastic convex optimization (MSCO) models on multi-commodity inventory problems and hydro-thermal power planning problems. The results show that our DR-MCO models could outperform MRCO and MSCO models when the data size is small.

math.OC

Moment Relaxations for Data-Driven Wasserstein Distributionally Robust Optimization

We propose moment relaxations for data-driven $p$-Wasserstein distributionally robust optimization ($p$-WDRO) problems that are defined by polynomials. The proposed moment relaxations admit Benders-type decomposition with parallel evaluation of the subgradients using each sample subproblem, which enables efficient solution for larger training sets. We then identify conditions on $p$ and the defining polynomial degrees such that the proposed $k$-th order moment relaxations preserve the asymptotic consistency of the original $p$-WDRO (i.e., the relaxation gap is bounded at most linearly by the Wasserstein radius). In particular, these conditions translate to effective bounds on $k$, which lead to polynomially sized semidefinite optimization formulations that are compatible with existing solvers. Numerical experiments on a box-constrained regression problem and a two-stage production problem are included to demonstrate the scalability and the effectiveness of the proposed moment relaxations.

math.OC

Semigroups of Integer Points in Convex Cones

We study the question whether the affine semigroup of integer points in a convex cone can be finitely generated up to symmetries of the cone. We establish general properties of finite generation up to symmetry, and then concentrate on the case of irrational polyhedral cones.

math.NT

Integer Points in Arbitrary Convex Cones: The Case of the PSD and SOC Cones

We investigate the semigroup of integer points inside a convex cone. We extend classical results in integer linear programming to integer conic programming. We show that the semigroup associated with nonpolyhedral cones can sometimes have a notion of finite generating set. We show this is true for the cone of positive semidefinite matrices (PSD) and the second-order cone (SOC). Both cones have a finite generating set of integer points, similar in spirit to Hilbert bases, under the action of a finitely generated group. We also extend notions of total dual integrality, Gomory-Chvátal closure, and Carathéodory rank to integer points in arbitrary cones.

math.OC

Spurious local minima in nonconvex sum-of-squares optimization

We study spurious second-order stationary points and local minima in a nonconvex low-rank formulation of sum-of-squares optimization on a real variety $X$. We reformulate the problem of finding a spurious local minimum in terms of syzygies of the underlying linear series, and also bring in topological tools to study this problem. When the variety $X$ is of minimal degree, there exist spurious second-order stationary points if and only if both the dimension and the codimension of the variety are greater than one, answering a question by Legat, Yuan, and Parrilo. Moreover, for surfaces of minimal degree, we provide sufficient conditions to exclude points from being spurious local minima. In particular, all second-order stationary points associated with infinite Gram matrices on the Veronese surface, corresponding to ternary quartics, lie on the boundary and can be written as a binary quartic, up to a linear change of coordinates, complementing work by Scheiderer on decompositions of ternary quartics as a sum of three squares. For general varieties of higher degree, we give examples and characterizations of spurious second-order stationary points in the interior, together with a restricted path algorithm that avoids such points with controlled step sizes, and numerical experiment results illustrating the empirical successes on plane cubic curves and Veronese varieties.

math.OC

A non-intrusive machine learning framework for debiasing long-time coarse resolution climate simulations and quantifying rare events statistics

Due to the rapidly changing climate, the frequency and severity of extreme weather is expected to increase over the coming decades. As fully-resolved climate simulations remain computationally intractable, policy makers must rely on coarse-models to quantify risk for extremes. However, coarse models suffer from inherent bias due to the ignored "sub-grid" scales. We propose a framework to non-intrusively debias coarse-resolution climate predictions using neural-network (NN) correction operators. Previous efforts have attempted to train such operators using loss functions that match statistics. However, this approach falls short with events that have longer return period than that of the training data, since the reference statistics have not converged. Here, the scope is to formulate a learning method that allows for correction of dynamics and quantification of extreme events with longer return period than the training data. The key obstacle is the chaotic nature of the underlying dynamics. To overcome this challenge, we introduce a dynamical systems approach where the correction operator is trained using reference data and a coarse model simulation nudged towards that reference. The method is demonstrated on debiasing an under-resolved quasi-geostrophic model and the Energy Exascale Earth System Model (E3SM). For the former, our method enables the quantification of events that have return period two orders longer than the training data. For the latter, when trained on 8 years of ERA5 data, our approach is able to correct the coarse E3SM output to closely reflect the 36-year ERA5 statistics for all prognostic variables and significantly reduce their spatial biases.

physics.ao-ph

An ADMM-based Distributed Optimization Method for Solving Security-Constrained AC Optimal Power Flow

In this paper, we study efficient and robust computational methods for solving the security-constrained alternating current optimal power flow (SC-ACOPF) problem, a two-stage nonlinear optimization problem with disjunctive constraints, that is central to the operation of electric power grids. The first-stage problem in SC-ACOPF determines the operation of the power grid in normal condition, while the second-stage problem responds to various contingencies of losing generators, transmission lines, and transformers. The two stages are coupled through disjunctive constraints, which model generators' active and reactive power output changes responding to system-wide active power imbalance and voltage deviations after contingencies. Real-world SC-ACOPF problems may involve power grids with more than 30k buses and 22k contingencies and need to be solved within 10-45 minutes to get a base case solution with high feasibility and reasonably good generation cost. We develop a comprehensive algorithmic framework to solve SC-ACOPF that meets the challenge of speed, solution quality, and computation robustness. In particular, we develop a smoothing technique to approximate disjunctive constraints into a smooth structure which can be handled by interior-point solvers; we design a distributed optimization algorithm to efficiently generate first-stage solutions; we propose a screening procedure to prioritize contingencies; and finally, we develop a reliable and parallel architecture that integrates all algorithmic components. Extensive tests on industry-scale systems demonstrate the superior performance of the proposed algorithms.

math.OC

Statistics of extreme events in coarse-scale climate simulations via machine learning correction operators trained on nudged datasets

This work presents a systematic framework for improving the predictions of statistical quantities for turbulent systems, with a focus on correcting climate simulations obtained by coarse-scale models. While high resolution simulations or reanalysis data are available, they cannot be directly used as training datasets to machine learn a correction for the coarse-scale climate model outputs, since chaotic divergence, inherent in the climate dynamics, makes datasets from different resolutions incompatible. To overcome this fundamental limitation we employ coarse-resolution model simulations nudged towards high quality climate realizations, here in the form of ERA5 reanalysis data. The nudging term is sufficiently small to not pollute the coarse-scale dynamics over short time scales, but also sufficiently large to keep the coarse-scale simulations close to the ERA5 trajectory over larger time scales. The result is a compatible pair of the ERA5 trajectory and the weakly nudged coarse-resolution E3SM output that is used as input training data to machine learn a correction operator. Once training is complete, we perform free-running coarse-scale E3SM simulations without nudging and use those as input to the machine-learned correction operator to obtain high-quality (corrected) outputs. The model is applied to atmospheric climate data with the purpose of predicting global and local statistics of various quantities of a time-period of a decade. Using datasets that are not employed for training, we demonstrate that the produced datasets from the ML-corrected coarse E3SM model have statistical properties that closely resemble the observations. Furthermore, the corrected coarse-scale E3SM output for the frequency of occurrence of extreme events, such as tropical cyclones and atmospheric rivers are presented. We present thorough comparisons and discuss limitations of the approach.

physics.ao-ph

Learning bias corrections for climate models using deep neural operators

Numerical simulation for climate modeling resolving all important scales is a computationally taxing process. Therefore, to circumvent this issue a low resolution simulation is performed, which is subsequently corrected for bias using reanalyzed data (ERA5), known as nudging correction. The existing implementation for nudging correction uses a relaxation based method for the algebraic difference between low resolution and ERA5 data. In this study, we replace the bias correction process with a surrogate model based on the Deep Operator Network (DeepONet). DeepONet (Deep Operator Neural Network) learns the mapping from the state before nudging (a functional) to the nudging tendency (another functional). The nudging tendency is a very high dimensional data albeit having many low energy modes. Therefore, the DeepoNet is combined with a convolution based auto-encoder-decoder (AED) architecture in order to learn the nudging tendency in a lower dimensional latent space efficiently. The accuracy of the DeepONet model is tested against the nudging tendency obtained from the E3SMv2 (Energy Exascale Earth System Model) and shows good agreement. The overarching goal of this work is to deploy the DeepONet model in an online setting and replace the nudging module in the E3SM loop for better efficiency and accuracy.

physics.ao-ph

Recent Developments in Security-Constrained AC Optimal Power Flow: Overview of Challenge 1 in the ARPA-E Grid Optimization Competition

The optimal power flow problem is central to many tasks in the design and operation of electric power grids. This problem seeks the minimum cost operating point for an electric power grid while satisfying both engineering requirements and physical laws describing how power flows through the electric network. By additionally considering the possibility of component failures and using an accurate AC power flow model of the electric network, the security-constrained AC optimal power flow (SC-AC-OPF) problem is of paramount practical relevance. To assess recent progress in solution algorithms for SC-AC-OPF problems and spur new innovations, the U.S. Department of Energy's Advanced Research Projects Agency--Energy (ARPA-E) organized Challenge 1 of the Grid Optimization (GO) competition. This paper describes the SC-AC-OPF problem formulation used in the competition, overviews historical developments and the state of the art in SC-AC-OPF algorithms, discusses the competition, and summarizes the algorithms used by the top three teams in Challenge 1 of the GO Competition (Teams gollnlp, GO-SNIP, and GMI-GO).

math.OC

Stochastic Dual Dynamic Programming for Multistage Stochastic Mixed-Integer Nonlinear Optimization

In this paper, we study multistage stochastic mixed-integer nonlinear programs (MS-MINLP). This general class of problems encompasses, as important special cases, multistage stochastic convex optimization with non-Lipschitzian value functions and multistage stochastic mixed-integer linear optimization. We develop stochastic dual dynamic programming (SDDP) type algorithms with nested decomposition, deterministic sampling, and stochastic sampling. The key ingredient is a new type of cuts based on generalized conjugacy. Several interesting classes of MS-MINLP are identified, where the new algorithms are guaranteed to obtain the global optimum without the assumption of complete recourse. This significantly generalizes the classic SDDP algorithms. We also characterize the iteration complexity of the proposed algorithms. In particular, for a $(T+1)$-stage stochastic MINLP with $d$-dimensional state spaces, to obtain an $ε$-optimal root node solution, we prove that the number of iterations of the proposed deterministic sampling algorithm is upper bounded by $\mathcal{O}((\frac{2T}ε)^d)$, and is lower bounded by $\mathcal{O}((\frac{T}{4ε})^d)$ for the general case or by $\mathcal{O}((\frac{T}{8ε})^{d/2-1})$ for the convex case. This shows that the obtained complexity bounds are rather sharp. It also reveals that the iteration complexity depends polynomially on the number of stages. We further show that the iteration complexity depends linearly on $T$, if all the state spaces are finite sets, or if we seek a $(Tε)$-optimal solution when the state spaces are infinite sets, i.e. allowing the optimality gap to scale with $T$. To the best of our knowledge, this is the first work that reports global optimization algorithms as well as iteration complexity results for solving such a large class of multistage stochastic programs.

math.OC

CondiDiag1.0: A flexible online diagnostic tool for conditional sampling and budget analysis in the E3SM atmosphere model (EAM)

Numerical models used in weather and climate prediction take into account a comprehensive set of atmospheric processes such as the resolved and unresolved fluid dynamics, radiative transfer, cloud and aerosol life cycles, and mass or energy exchanges with the Earth's surface. In order to identify model deficiencies and improve predictive skills, it is important to obtain process-level understanding of the interactions between different processes. Conditional sampling and budget analysis are powerful tools for process-oriented model evaluation, but they often require tedious ad hoc coding and large amounts of instantaneous model output, resulting in inefficient use of human and computing resources. This paper presents an online diagnostic tool that addresses this challenge by monitoring model variables in a generic manner as they evolve within the time integration cycle. The tool is convenient to use. It allows users to select sampling conditions and specify monitored variables at run time. Both the evolving values of the model variables and their increments caused by different atmospheric processes can be monitored and archived. Online calculation of vertical integrals is also supported. Multiple sampling conditions can be monitored in a single simulation in combination with unconditional sampling. The paper explains in detail the design and implementation of the tool in the Energy Exascale Earth System Model (E3SM) version 1. The usage is demonstrated through three examples: a global budget analysis of dust aerosol mass concentration, a composite analysis of sea salt emission and its dependency on surface wind speed, and a conditionally sampled relative humidity budget. The tool is expected to be easily portable to closely related atmospheric models that use the same or similar data structures and time integration methods.

physics.ao-ph

Quantifying and attributing time step sensitivities in present-day climate simulations conducted with EAMv1

This study assesses the relative importance of time integration error in present-day climate simulations conducted with the atmosphere component of the Energy Exascale Earth System Model version 1 (EAMv1) at 1-degree horizontal resolution. We show that a factor-of-6 reduction of time step size in all major parts of the model leads to significant changes in the long-term mean climate. These changes imply that the reduction of temporal truncation errors leads to a notable although unsurprising degradation of agreement between the simulated and observed present-day climate; the model would require retuning to regain optimal climate fidelity in the absence of those truncation errors. A coarse-grained attribution of the time step sensitivities is carried out by separately shortening time steps used in various components of EAM or by revising the numerical coupling between some processes. The results provide useful clues to help better understand the root causes of time step sensitivities in EAM. The experimentation strategy used here can also provide a pathway for other models to identify and reduce time integration errors.

physics.ao-ph

On Distributionally Robust Multistage Convex Optimization: New Algorithms and Complexity Analysis

This paper presents an algorithmic study and complexity analysis for solving distributionally robust multistage convex optimization (DR-MCO) problems. Our main contribution is a novel nonconsecutive dual dynamic programming (NDDP) algorithm which explores different stages in an adaptive fashion. In contrast with the usual consecutive dual dynamic programming (CDDP) algorithm, we show that NDDP reduces the subproblem complexity from quadratic to linear dependency on the number of stages. Two different DR-MCO examples are also presented to show the efficiency and effectiveness of the proposed NDDP algorithm.

math.OC