SearcharxivSearch

arXiv subjects

Donghyun You

Publications and source records attributed to Donghyun You.

17 recordsLinked to original sources

A Highly Scalable TDMA for GPUs and Its Application to Flow Solver Optimization

A tridiagonal matrix algorithm (TDMA), Pipelined-TDMA, is developed for multi-GPU systems to resolve the scalability bottlenecks caused by the sequential structure of conventional divide-and-conquer TDMA. The proposed method pipelines multiple tridiagonal systems, overlapping communication with computation and executing GPU kernels concurrently to hide non-scalable stages behind scalable compute stages. To maximize performance, the batch size is optimized to strike a balance between GPU occupancy and pipeline efficiency: larger batches improve throughput for solving tridiagonal systems, while excessively large batches reduce pipeline utilization. Performance evaluations on up to 64 NVIDIA A100 GPUs using a one-dimensional (1D) slab-type domain decomposition confirm that, except for the terminal phase of the pipeline, the proposed method successfully hides most of the non-scalable execution time-specifically inter-GPU communication and low-occupancy computation. The solver achieves ideal weak scaling up to 64 GPUs with one billion grid cells per GPU and reaches 74.7 percent of ideal performance in strong scaling tests for a 4-billion-cell problem, relative to a 4-GPU baseline. The optimized TDMA is integrated into an ADI-based fractional-step method to remove the scalability bottleneck in the Poisson solver of the flow solver (Ha et al., 2021). In a 9-billion-cell simulation on 64 GPUs, the TDMA component in the Poisson solver is accelerated by 4.37x, contributing to a 1.31x overall speedup of the complete flow solver.

physics.comp-ph

Optimal mesh generation for a non-iterative grid-converged solution of flow through a blade passage using deep reinforcement learning

An automatic mesh generation method for optimal computational fluid dynamics (CFD) analysis of a blade passage is developed using deep reinforcement learning (DRL). Unlike conventional automation techniques, which require repetitive tuning of meshing parameters for each new geometry and flow condition, the method developed herein trains a mesh generator to determine optimal parameters across varying configurations in a non-iterative manner. Initially, parameters controlling mesh shape are optimized to maximize geometric mesh quality, as measured by the ratio of determinants of Jacobian matrices and skewness. Subsequently, resolution-controlling parameters are optimized by incorporating CFD results. Multi-agent reinforcement learning is employed, enabling 256 agents to construct meshes and perform CFD analyses across randomly assigned flow configurations in parallel, aiming for maximum simulation accuracy and computational efficiency within a multi-objective optimization framework. After training, the mesh generator is capable of producing meshes that yield converged solutions at desired computational costs for new configurations in a single simulation, thereby eliminating the need for iterative CFD procedures for grid convergence. The robustness and effectiveness of the method are investigated across various blade passage configurations, accommodating a range of blade geometries, including high-pressure and low-pressure turbine blades, axial compressor blades, and impulse rotor blades. Furthermore, the method is capable of identifying the optimal mesh resolution for diverse flow conditions, including complex phenomena like boundary layers, shock waves, and flow separation. The optimality is confirmed by comparing the accuracy and the efficiency achieved in a single attempt with those from the conventional iterative optimization method.

physics.flu-dyn

Non-iterative generation of an optimal mesh for a blade passage using deep reinforcement learning

A method using deep reinforcement learning (DRL) to non-iteratively generate an optimal mesh for an arbitrary blade passage is developed. Despite automation in mesh generation using either an empirical approach or an optimization algorithm, repeated tuning of meshing parameters is still required for a new geometry. The method developed herein employs a DRL-based multi-condition optimization technique to define optimal meshing parameters as a function of the blade geometry, attaining automation, minimization of human intervention, and computational efficiency. The meshing parameters are optimized by training an elliptic mesh generator which generates a structured mesh for a blade passage with an arbitrary blade geometry. During each episode of the DRL process, the mesh generator is trained to produce an optimal mesh for a randomly selected blade passage by updating the meshing parameters until the mesh quality, as measured by the ratio of determinants of the Jacobian matrices and the skewness, reaches the highest level. Once the training is completed, the mesh generator create an optimal mesh for a new arbitrary blade passage in a single try without an repetitive process for the parameter tuning for mesh generation from the scratch. The effectiveness and robustness of the proposed method are demonstrated through the generation of meshes for various blade passages.

cs.LG

Neural-network-based mixed subgrid-scale model for turbulent flow

An artificial neural-network-based subgrid-scale model using the resolved stress, which is capable of predicting untrained decaying isotropic turbulence, is developed. Providing the grid-scale strain-rate tensor alone as input leads the model to predict a subgrid-scale stress tensor aligns with the strain-rate tensor, and the model performs similar to the dynamic Smagorinsky model. On the other hand, providing the resolved stress tensor as input in addition to the strain-rate tensor is found to significantly improve the model in terms of the energy spectra and probability density function of subgrid-scale dissipation. In an attempt to apply the neural-network-based model trained for forced homogeneous isotropic turbulence to decaying homogeneous isotropic turbulence, special attention is given to the normalisation of the input and output tensors. It is found that successful generalisation of the model to turbulence at various untrained conditions is possible if the distributions of the normalised inputs and outputs of the neural-network remain unchanged as Reynolds numbers and grid resolution of the turbulence vary. In a posteriori tests of the forced and the decaying homogeneous isotropic turbulence, the developed neural-network model is found to predict turbulence statistics more accurately and to be computationally more efficient than the conventional dynamic models.

physics.flu-dyn

Particle swarm optimization of a wind farm layout with active control of turbine yaws

Active yaw control (AYC) of wind turbines has been widely applied to increase the annual energy production (AEP) of a wind farm. AYC efficiency depends on the wind direction and the wind farm layout because an AYC method utilizes wake deflection by yawing wind turbines. Conventional optimization of a wind farm layout assumed that the swept areas of all wind turbines are aligned perpendicular to the wind direction, thereby allowing non-optimal utilization of an AYC method. Higher AEP can be obtained by joint optimization which considers an AYC method in the layout design stage. Joint optimization of the farm layout and AYC has been difficult due to the non-convexity of the problem and the computational inefficiency. In the present study, a particle swarm optimization based method is developed for joint optimization. The layout is optimized with simultaneous consideration for yaw angles for all wind velocities to obtain a globally optimal layout. A number of random initial particles consisting of the layout and yaw angles of wind turbines reduce the initial layout dependency on the optimized layout. To deal with the challenge of large-scale optimization, the adaptive granularity learning distributed particle swarm optimization algorithm is implemented. The improvement in AEP when using a jointly optimized layout compared to a conventionally optimized layout in a real wind farm is demonstrated using the present method.

math.OC

A realizable second-order advection method with variable flux limiters for moment transport equations

A second-order total variation diminishing (TVD) method with variable flux limiters is proposed to overcome the non-realizability issue, which has been one of major obstacles in applying the conventional second-order TVD schemes to the moment transport equations. In the present method, a realizable moment set at a cell face is reconstructed by allowing the flexible selection of the flux limiter values within the second-order TVD region. Necessary conditions for the variable flux limiter scheme to simultaneously satisfy the realizability and the second-order TVD property for the third-order moment set are proposed. The strategy for satisfying the second-order TVD property is conditionally extended to the fourth- and fifth-order moments. The proposed method is verified and compared with other high-order realizable schemes in one- and two-dimensional configurations, and is found to preserve the realizability of moments while satisfying the high-order TVD property for the third-order moment set and conditionally for the fourth- and fifth-order moments.

physics.flu-dyn

Control of a fly-mimicking flyer in complex flow using deep reinforcement learning

An integrated framework of computational fluid-structural dynamics (CFD-CSD) and deep reinforcement learning (deep-RL) is developed for control of a fly-scale flexible-winged flyer in complex flow. Dynamics of the flyer in complex flow is highly unsteady and nonlinear, which makes modeling the dynamics challenging. Thus, conventional control methodologies, where the dynamics is modeled, are insufficient for regulating such complicated dynamics. Therefore, in the present study, the integrated framework, in which the whole governing equations for fluid and structure are solved, is proposed to generate a control policy for the flyer. For the deep-RL to successfully learn the control policy, accurate and ample data of the dynamics are required. However, satisfying both the quality and quantity of the data on the intricate dynamics is extremely difficult since, in general, more accurate data are more costly. In the present study, two strategies are proposed to deal with the dilemma. To obtain accurate data, the CFD-CSD is adopted for precisely predicting the dynamics. To gain ample data, a novel data reproduction method is devised, where the obtained data are replicated for various situations while conserving the dynamics. With those data, the framework learns the control policy in various flow conditions and the learned policy is shown to have remarkable performance in controlling the flyer in complex flow fields.

cs.LG

Multi-condition multi-objective optimization using deep reinforcement learning

A multi-condition multi-objective optimization method that can find Pareto front over a defined condition space is developed for the first time using deep reinforcement learning. Unlike the conventional methods which perform optimization at a single condition, the present method learns the correlations between conditions and optimal solutions. The exclusive capability of the developed method is examined in the solutions of a novel modified Kursawe benchmark problem and an airfoil shape optimization problem which include nonlinear characteristics which are difficult to resolve using conventional optimization methods. Pareto front with high resolution over a defined condition space is successfully determined in each problem. Compared with multiple operations of a single-condition optimization method for multiple conditions, the present multi-condition optimization method based on deep reinforcement learning shows a greatly accelerated search of Pareto front by reducing the number of required function evaluations. An analysis of aerodynamics performance of airfoils with optimally designed shapes confirms that multi-condition optimization is indispensable to avoid significant degradation of target performance for varying flow conditions.

cs.LG

Numerical modeling of bubble-particle interaction in a volume-of-fluid framework

A numerical method is presented to simulate gas-liquid-solid flows with bubble-particle interaction, including particle collision, sliding, and attachment. Gas-liquid flows are simulated in an Eulerian framework using a volume-of-fluid method. Particle motions are predicted in a Lagrangian framework. Algorithms that are used to detect collision and determine the sliding or attachment of the particle are developed. An effective bubble is introduced to model these bubble-particle interaction. The proposed numerical method is validated through experimental cases that entail the rising of a single bubble with particles. Collision and attachment probabilities obtained from the simulation are compared to model and experimental results based on bubble diameters, particle diameters, and contact angles. The particle trajectories near the bubble are presented to show differences with and without the proposed bubble-particle interaction model. The sliding and attachment of the colliding particle are observed using this model.

physics.flu-dyn

Mechanisms of a Convolutional Neural Network for Learning Three-dimensional Unsteady Wake Flow

Convolutional neural networks (CNNs) have recently been applied to predict or model fluid dynamics. However, mechanisms of CNNs for learning fluid dynamics are still not well understood, while such understanding is highly necessary to optimize the network or to reduce trial-and-errors during the network optmization. In the present study, a CNN to predict future three-dimensional unsteady wake flow using flow fields in the past occasions is developed. Mechanisms of the developed CNN for prediction of wake flow behind a circular cylinder are investigated in two flow regimes: the three-dimensional wake transition regime and the shear-layer transition regime. Feature maps in the CNN are visualized to compare flow structures which are extracted by the CNN from flow at the two flow regimes. In both flow regimes, feature maps are found to extract similar sets of flow structures such as braid shear-layers and shedding vortices. A Fourier analysis is conducted to investigate mechanisms of the CNN for predicting wake flow in flow regimes with different wave number characteristics. It is found that a convolution layer in the CNN integrates and transports wave number information from flow to predict the dynamics. Characteristics of the CNN for transporting input information including time histories of flow variables is analyzed by assessing contributions of each flow variable and time history to feature maps in the CNN. Structural similarities between feature maps in the CNN are calculated to reveal the number of feature maps that contain similar flow structures. By reducing the number of feature maps that contain similar flow structures, it is also able to successfully reduce the number of parameters to learn in the CNN by 85\% without affecting prediction performances.

physics.flu-dyn

Data-driven prediction of unsteady flow fields over a circular cylinder using deep learning

Unsteady flow fields over a circular cylinder are trained and predicted using four different deep learning networks: convolutional neural networks with and without consideration of conservation laws, generative adversarial networks with and without consideration of conservation laws. Flow fields at future occasions are predicted based on information of flow fields at previous occasions. Deep learning networks are trained first using flow fields at Reynolds numbers of 100, 200, 300, and 400, while flow fields at Reynolds numbers of 500 and 3000 are predicted using the trained deep learning networks. Physical loss functions are proposed to explicitly impose information of conservation of mass and momentum to deep learning networks. An adversarial training is applied to extract features of flow dynamics in an unsupervised manner. Effects of the proposed physical loss functions, adversarial training, and network sizes on the prediction accuracy are analyzed. Predicted flow fields using deep learning networks are in favorable agreement with flow fields computed by numerical simulations.

physics.flu-dyn

Prediction of typhoon tracks using a generative adversarial network with observational and meteorological data

Tracks of typhoons are predicted using a generative adversarial network (GAN) with observational data in form of satellite images and meteorological data from a reanalysis database. Time series of images of typhoons which occurred in the Korean Peninsula in the past are used to train the neural network. The trained GAN is employed to produce a 6-hour-advance track of a typhoon for which the GAN was not trained. The predicted image favorably identifies the future location of the typhoon center as well as the deformed cloud structures. The errors between predicted and real typhoon centers are measured quantitatively in kilometers. 65.5 % of all typhoon center predictions have an error of less than 80 km, 31.5 % lie within a range of 80 - 120 km and the remaining 3.0 % are above 120 km. The overall error is 67.2 km, compared to 95.6 km when only observational data are used as input. The cloud structure prediction is evaluated qualitatively. It is shown that the GAN is able to predict trends in cloud motion. It is found that adding physically meaningful meteorological data to satellite images improves the sharpness of predicted images.

physics.ao-ph

A scalable multi-GPU method for semi-implicit fractional-step integration of incompressible Navier-Stokes equations

A new flow solver scalable on multiple Graphics Processing Units (GPUs) for direct numerical simulation of wall-bounded incompressible flow is presented. This solver utilizes a previously reported work (J. Comp. Physics, vol. 352 (2018), pp.246-264) which proposes a semi-implicit fractional-step method on a single GPU. Extension of this work to accommodate multiple GPUs becomes inefficient when global transpose is used in the Alternating Direction Implicit (ADI) and Fourier-transform-based direct methods. A new strategy for designing an efficient multi-GPU solver is described to completely remove global transpose and achieve high scalability. Parallel Diagonal Dominant (PDD) and Parallel Partition (PPT) methods are implemented for GPUs to obtain good scaling and preserve accuracy. An overall efficiency of 0.89 is shown. Turbulent flat-plate boundary layer is simulated on 607M grid points using 4 Tesla P100 GPUs.

physics.comp-ph

Deep learning approach in multi-scale prediction of turbulent mixing-layer

Achievement of solutions in Navier-Stokes equation is one of challenging quests, especially for its closure problem. For achievement of particular solutions, there are variety of numerical simulations including Direct Numerical Simulation (DNS) or Large Eddy Simulation (LES). These methods analyze flow physics through efficient reduced-order modeling such as proper orthogonal decomposition or Koopman method, showing prominent fidelity in fluid dynamics. Generative adversarial network (GAN) is a reprint of neurons in brain as combinations of linear operations, using competition between generator and discriminator. Current paper propose deep learning network for prediction of small-scale movements with large-scale inspections only, using GAN. Therefore DNS result of three-dimensional mixing-layer was filtered blurring out the small-scaled structures, then is predicted of its detailed structures, utilizing Generative Adversarial Network (GAN). This enables multi-resolution analysis being asked to predict fine-resolution solution with only inspection of blurry one. Within the grid scale, current paper present deep learning approach of modeling small scale features in turbulent flow. The presented method is expected to have its novelty in utilization of unprocessed simulation data, achievement of 3D structures in prediction by processing 3D convolutions, and predicting precise solution with less computational costs.

physics.comp-ph

Typhoon track prediction using satellite images in a Generative Adversarial Network

Tracks of typhoons are predicted using satellite images as input for a Generative Adversarial Network (GAN). The satellite images have time gaps of 6 hours and are marked with a red square at the location of the typhoon center. The GAN uses images from the past to generate an image one time step ahead. The generated image shows the future location of the typhoon center, as well as the future cloud structures. The errors between predicted and real typhoon centers are measured quantitatively in kilometers. 42.4% of all typhoon center predictions have absolute errors of less than 80 km, 32.1% lie within a range of 80 - 120 km and the remaining 25.5% have accuracies above 120 km. The relative error sets the above mentioned absolute error in relation to the distance that has been traveled by a typhoon over the past 6 hours. High relative errors are found in three types of situations, when a typhoon moves on the open sea far away from land, when a typhoon changes its course suddenly and when a typhoon is about to hit the mainland. The cloud structure prediction is evaluated qualitatively. It is shown that the GAN is able to predict trends in cloud motion. In order to improve both, the typhoon center and cloud motion prediction, the present study suggests to add information about the sea surface temperature, surface pressure and velocity fields to the input data.

physics.ao-ph

Prediction of laminar vortex shedding over a cylinder using deep learning

Unsteady laminar vortex shedding over a circular cylinder is predicted using a deep learning technique, a generative adversarial network (GAN), with a particular emphasis on elucidating the potential of learning the solution of the Navier-Stokes equations. Numerical simulations at two different Reynolds numbers with different time-step sizes are conducted to produce training datasets of flow field variables. Unsteady flow fields in the future at a Reynolds number which is not in the training datasets are predicted using a GAN. Predicted flow fields are found to qualitatively and quantitatively agree well with flow fields calculated by numerical simulations. The present study suggests that a deep learning technique can be utilized for prediction of laminar wake flow in lieu of solving the Navier-Stokes equations.

physics.flu-dyn

An adaptive memory method for accurate and efficient computation of the Caputo fractional derivative

A fractional derivative is a temporally nonlocal operation which is computationally intensive due to inclusion of the accumulated contribution of function values at past times. In order to lessen the computational load while maintaining the accuracy of the fractional derivative, a novel numerical method for the Caputo fractional derivative is proposed. The present adaptive memory method significantly reduces the requirement for computational memory for storing function values at past time points and also significantly improves the accuracy by calculating convolution weights to function values at past time points which can be non-uniformly distributed in time. The superior accuracy of the present method to the accuracy of the previously reported methods is identified by deriving numerical errors analytically. The sub-diffusion process of a time-fractional diffusion equation and the creeping response of a fractional viscoelastic model are simulated to demonstrate the accuracy as well as the computational efficiency of the present method.

math.NA