SearcharxivSearch

arXiv subjects

Michel Schanen

Publications and source records attributed to Michel Schanen.

13 recordsLinked to original sources

ExaModels.jl: an Algebraic Modeling System for Nonlinear Programming on GPUs

Large-scale nonlinear programs almost always exhibit partially separable and repetitive structure, yet most existing algebraic modeling systems do not take advantage of it. A nonlinear optimization solver queries the objective, the constraints, and their derivatives at every iteration, so the speed of these evaluations bears directly on the overall solution time. We present ExaModels.jl, a Julia-based algebraic modeling system that exploits this structure to evaluate the objective, the constraints, and their derivatives in parallel. At its core is a single-instruction, multiple-data abstraction that represents a nonlinear program as a small number of algebraic patterns, each repeated over many data points. Because the patterns are visible at compile time, a specialized model and derivative evaluation kernel is compiled for each pattern. Applying that kernel independently across the data points maps naturally onto GPU parallelism and, with sufficiently many threads, yields O(1) evaluation time regardless of the number of data points. On the largest instances of the Luksan-Vlcek library, GPU execution speeds up sparse Hessian evaluation by 76x over single-threaded CPU evaluation, and by 30x on COPS and 7.3x on PGLIB-OPF.

math.OC

Condensed Interior-Point Methods for Scalable Nonlinear Programming on GPUs

This paper explores two variants of condensed-space interior-point methods designed for GPUs - HyKKT and LiftedKKT - by analyzing their numerical properties through error analysis and assessing their real-world performance via extensive numerical experiments with a fully GPU-resident software implementation. Traditional implementations of interior-point methods (IPMs) involve solving indefinite augmented KKT systems repeatedly by utilizing direct sparse solvers based on the LBL factorization with sophisticated numerical pivoting strategies. While this method achieves high performance and robustness on CPUs, the serial nature of numerical pivoting presents challenges for effective implementation on GPUs. Recently, multiple condensed-space IPM strategies have emerged to address this issue by transforming the KKT system into a symmetric positive-definite matrix, which is more suitable for factorization on GPUs. In this study, we demonstrate that although the condensed systems show increased ill-conditioning, the inherent structures of the condensed KKT system effectively counterbalance potential accuracy loss in the IPM. Furthermore, we provide numerical results that thoroughly assess the capabilities of a fully GPU-resident nonlinear programming software stack, comprising MadNLP (a filter line-search IPM solver), cuDSS (a direct sparse solver leveraging Cholesky factorization), and ExaModels (a modeling framework), by benchmarking their performance against the pglib-opf and CUTEst libraries. Our findings suggest that the GPU framework holds promise for solving highly sparse large-scale nonlinear programs, such as optimal power flow instances, except for diminished robustness and limited speedups for edge cases observed within CUTEst instances.

math.OC

Harnessing GPU Acceleration in Large-Scale Process Optimization

This paper presents a proof-of-concept workflow for equation-oriented process optimization that runs entirely on a GPU. Process optimization models often incorporate complex interconnected unit operations, dynamics, and uncertainties, resulting in large nonlinear programs that can be computationally demanding for conventional CPU-based solvers. Although emerging GPU-based solvers offer substantial computational benefits, their application to process optimization has been limited by the lack of GPU-compatible process modeling tools. We address this gap by prototyping the GPU-compatible process optimization models using an existing GPU-capable optimization software stack, including ExaModels (algebraic modeling system), MadNLP (optimization solver), and cuDSS (linear solver). ExaModels formulates the process optimization problem in a GPU-compatible way by exposing its repeated algebraic structure, while MadNLP and cuDSS solve the resulting nonlinear program on the GPU. This workflow is demonstrated on a CO2 absorber design problem under feed uncertainty, in which a shared column diameter is minimized subject to equilibrium and hydraulic constraints in all scenarios. For the largest case with 5,000 scenarios and 1.5 million variables, the GPU workflow achieves a speedup of approximately 21\times over a single-threaded CPU baseline using JuMP, Ipopt, and MA57.

math.OC

Scalable Multi-Period AC Optimal Power Flow Utilizing GPUs with High Memory Capacities

This paper demonstrates the scalability of open-source GPU-accelerated nonlinear programming (NLP) frameworks -- ExaModels.jl and MadNLP.jl -- for solving multi-period alternating current (AC) optimal power flow (OPF) problems on GPUs with high memory capacities (e.g., NVIDIA GH200 with 480 GB of unified memory). There has been a growing interest in solving multi-period AC OPF problems, as the increasingly fluctuating electricity market requires operation planning over multiple periods. These problems, formerly deemed intractable, are now becoming technologically feasible to solve thanks to the advent of high-memory GPU hardware and accelerated NLP tools. This study evaluates the capability of these tools to tackle previously unsolvable multi-period AC OPF instances. Our numerical experiments, run on an NVIDIA GH200, demonstrate that we can solve a multi-period OPF instance with more than 10 million variables up to $10^{-4}$ precision in less than 10 minutes. These results demonstrate the efficacy of the GPU-accelerated NLP frameworks for the solution of extreme-scale multi-period OPF. We provide ExaModelsPower.jl, an open-source modeling tool for multi-period AC OPF models for GPUs.

math.OC

GPU-Accelerated Sequential Quadratic Programming Algorithm for Solving ACOPF

Sequential quadratic programming (SQP) is widely used in solving nonlinear optimization problem, with advantages of warm-starting solutions, as well as finding high-accurate solution and converging quadratically using second-order information, such as the Hessian matrix. In this study we develop a scalable SQP algorithm for solving the alternate current optimal power flow problem (ACOPF), leveraging the parallel computing capabilities of graphics processing units (GPUs). Our methodology incorporates the alternating direction method of multipliers (ADMM) to initialize and decompose the quadratic programming subproblems within each SQP iteration into independent small subproblems for each electric grid component. We have implemented the proposed SQP algorithm using our portable Julia package ExaAdmm.jl, which solves the ADMM subproblems in parallel on all major GPU architectures. For numerical experiments, we compared three solution approaches: (i) the SQP algorithm with a GPU-based ADMM subproblem solver, (ii) a CPU-based ADMM solver, and (iii) the QP solver Ipopt (the state-of-the-art interior point solver) and observed that for larger instances our GPU-based SQP solver efficiently leverages the GPU many-core architecture, dramatically reducing the solution time.

math.OC

Parallel Interior-Point Solver for Block-Structured Nonlinear Programs on SIMD/GPU Architectures

We investigate how to port the standard interior-point method to new exascale architectures for block-structured nonlinear programs with state equations. Computationally, we decompose the interior-point algorithm into two successive operations: the evaluation of the derivatives and the solution of the associated Karush-Kuhn-Tucker (KKT) linear system. Our method accelerates both operations using two levels of parallelism. First, we distribute the computations on multiple processes using coarse parallelism. Second, each process uses a SIMD/GPU accelerator locally to accelerate the operations using fine-grained parallelism. The KKT system is reduced by eliminating the inequalities and the state variables from the corresponding equations, to a dense matrix encoding the sensitivities of the problem's degrees of freedom, drastically minimizing the memory exchange. We demonstrate the method's capability on the supercomputer Polaris, a testbed for the future exascale Aurora system. Each node is equipped with four GPUs, a setup amenable to our two-level approach. Our experiments on the stochastic optimal power flow problem show that the method can achieve a 50x speed-up compared to the state-of-the-art method.

math.OC

Fitting Matérn Smoothness Parameters Using Automatic Differentiation

The Matérn covariance function is ubiquitous in the application of Gaussian processes to spatial statistics and beyond. Perhaps the most important reason for this is that the smoothness parameter $ν$ gives complete control over the mean-square differentiability of the process, which has significant implications for the behavior of estimated quantities such as interpolants and forecasts. Unfortunately, derivatives of the Matérn covariance function with respect to $ν$ require derivatives of the modified second-kind Bessel function $\mathcal{K}_ν$ with respect to $ν$. While closed form expressions of these derivatives do exist, they are prohibitively difficult and expensive to compute. For this reason, many software packages require fixing $ν$ as opposed to estimating it, and all existing software packages that attempt to offer the functionality of estimating $ν$ use finite difference estimates for $\partial_ν\mathcal{K}_ν$. In this work, we introduce a new implementation of $\mathcal{K}_ν$ that has been designed to provide derivatives via automatic differentiation (AD), and whose resulting derivatives are significantly faster and more accurate than those computed using finite differences. We provide comprehensive testing for both speed and accuracy and show that our AD solution can be used to build accurate Hessian matrices for second-order maximum likelihood estimation in settings where Hessians built with finite difference approximations completely fail.

stat.CO

Condensed interior-point methods: porting reduced-space approaches on GPU hardware

The interior-point method (IPM) has become the workhorse method for nonlinear programming. The performance of IPM is directly related to the linear solver employed to factorize the Karush--Kuhn--Tucker (KKT) system at each iteration of the algorithm. When solving large-scale nonlinear problems, state-of-the art IPM solvers rely on efficient sparse linear solvers to solve the KKT system. Instead, we propose a novel reduced-space IPM algorithm that condenses the KKT system into a dense matrix whose size is proportional to the number of degrees of freedom in the problem. Depending on where the reduction occurs we derive two variants of the reduced-space method: linearize-then-reduce and reduce-then-linearize. We adapt their workflow so that the vast majority of computations are accelerated on GPUs. We provide extensive numerical results on the optimal power flow problem, comparing our GPU-accelerated reduced space IPM with Knitro and a hybrid full space IPM algorithm. By evaluating the derivatives on the GPU and solving the KKT system on the CPU, the hybrid solution is already significantly faster than the CPU-only solutions. The two reduced-space algorithms go one step further by solving the KKT system entirely on the GPU. As expected, the performance of the two reduction algorithms depends intrinsically on the number of available degrees of freedom: their performance is poor when the problem has many degrees of freedom, but the two algorithms are up to 3 times faster than Knitro as soon as the relative number of degrees of freedom becomes smaller.

math.OC

On automatic differentiation for the Matérn covariance

To target challenges in differentiable optimization we analyze and propose strategies for derivatives of the Matérn kernel with respect to the smoothness parameter. This problem is of high interest in Gaussian processes modelling due to the lack of robust derivatives of the modified Bessel function of second kind with respect to order. In the current work we focus on newly identified series expansions for the modified Bessel function of second kind valid for complex orders. Using these expansions we obtain highly accurate results using the complex step method. Furthermore, we show that the evaluations using the recommended expansions are also more efficient than finite differences.

math.NA

Batched Second-Order Adjoint Sensitivity for Reduced Space Methods

This paper presents an efficient method for extracting the second-order sensitivities from a system of implicit nonlinear equations on upcoming graphical processing units (GPU) dominated computer systems. We design a custom automatic differentiation (AutoDiff) backend that targets highly parallel architectures by extracting the second-order information in batch. When the nonlinear equations are associated to a reduced space optimization problem, we leverage the parallel reverse-mode accumulation in a batched adjoint-adjoint algorithm to compute efficiently the reduced Hessian of the problem. We apply the method to extract the reduced Hessian associated to the balance equations of a power network, and show on the largest instances that a parallel GPU implementation is 30 times faster than a sequential CPU reference based on UMFPACK.

cs.MS

A Globally Convergent Distributed Jacobi Scheme for Block-Structured Nonconvex Constrained Optimization Problems

Motivated by the increasing availability of high-performance parallel computing, we design a distributed parallel algorithm for linearly-coupled block-structured nonconvex constrained optimization problems. Our algorithm performs Jacobi-type proximal updates of the augmented Lagrangian function, requiring only local solutions of separable block nonlinear programming (NLP) problems. We provide a cheap and explicitly computable Lyapunov function that allows us to establish global and local sublinear convergence of our algorithm, its iteration complexity, as well as simple, practical and theoretically convergent rules for automatically tuning its parameters. This in contrast to existing algorithms for nonconvex constrained optimization based on the alternating direction method of multipliers that rely on at least one of the following: Gauss-Seidel or sequential updates, global solutions of NLP problems, non-computable Lyapunov functions, and hand-tuning of parameters. Numerical experiments showcase its advantages for large-scale problems, including the multi-period optimization of a 9000-bus AC optimal power flow test case over 168 time periods, solved on the Summit supercomputer using an open-source Julia code.

math.OC

A Feasible Reduced Space Method for Real-Time Optimal Power Flow

We propose a novel feasible-path algorithm to solve the optimal power flow (OPF) problem for real-time use cases. The method augments the seminal work of Dommel and Tinney with second-order derivatives to work directly in the reduced space induced by the power flow equations. In the reduced space, the optimization problem includes only inequality constraints corresponding to the operational constraints. While the reduced formulation directly enforces the physical constraints, the operational constraints are softly enforced through Augmented Lagrangian penalty terms. In contrast to interior-point algorithms (state-of-the art for solving OPF), our algorithm maintains feasibility at each iteration, which makes it suitable for real-time application. By exploiting accelerator hardware (Graphic Processing Units) to compute the reduced Hessian, we show that the second-order method is numerically tractable and is effective to solve both static and real-time OPF problems.

math.OC

On the Strong Scaling of the Spectral Element Solver Nek5000 on Petascale Systems

The present work is targeted at performing a strong scaling study of the high-order spectral element fluid dynamics solver Nek5000. Prior studies indicated a recommendable metric for strong scalability from a theoretical viewpoint, which we test here extensively on three parallel machines with different performance characteristics and interconnect networks, namely Mira (IBM Blue Gene/Q), Beskow (Cray XC40) and Titan (Cray XK7). The test cases considered for the simulations correspond to a turbulent flow in a straight pipe at four different friction Reynolds numbers $Re_τ$ = 180, 360, 550 and 1000. Considering the linear model for parallel communication we quantify the machine characteristics in order to better assess the scaling behaviors of the code. Subsequently sampling and profiling tools are used to measure the computation and communication times over a large range of compute cores. We also study the effect of the two coarse grid solvers XXT and AMG on the computational time. Super-linear scaling due to a reduction in cache misses is observed on each computer. The strong scaling limit is attained for roughly 5000 - 10,000 degrees of freedom per core on Mira, 30,000 - 50,0000 on Beskow, with only a small impact of the problem size for both machines, and ranges between 10,000 and 220,000 depending on the problem size on Titan. This work aims at being a reference for Nek5000 users and also serves as a basis for potential issues to address as the community heads towards exascale supercomputers.

cs.DC