SearcharxivSearch

arXiv subjects

Christopher W. Tessum

Publications and source records attributed to Christopher W. Tessum.

7 recordsLinked to original sources

Acceleration of horizontal numerical advection for atmospheric modeling through surrogate modeling with temporal coarse-graining

Machine-learned surrogate modeling of advection may accelerate geoscientific models, but existing approaches have either achieved limited speedup or have sacrificed spatial resolution compared to the model they are trained to emulate. We developed a machine-learned solver that speeds up advection simulations without sacrificing spatial resolution through the use of temporal coarse-graining, where the model is trained to take larger integration steps than dictated by the Courant-Friedrich-Lewy (CFL) condition. Our solver framework includes a convolutional neural network that takes concentrations and CFL numbers as inputs and outputs mass flux. Our solvers emulate 10-day ground-level horizontal advection simulations with r$^2$ values against the baseline ranging from 0.60--0.98 with temporal coarsening factors of 4 to 32 times the baseline integration time step. Speed increases and accuracy decreases with increased coarsening, with $r^2 = 0.24$ in accuracy lost for every factor of 10 gained in speed, reaching a maximum 92$\times$ speedup while maintaining $r^2 = 0.60$. We deliberately trained our solvers only on January ground-level wind data to examine their ability to generalize across seasons and vertical heights. The 4$\times$-coarsened learned solver successfully reproduces simulations over 72 vertical levels. The 8$\times$--16$\times$ solvers (but not 32$\times$) emulate most vertical levels. The learned solvers also generalize well across seasons, except for instabilities in June and October. With additional fine-tuning, these learned solvers could be appropriate for operational use where trading accuracy for speed could be advantageous, such as in screening tools, in ensemble simulations, or with data assimilation.

physics.ao-ph

AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization

Image labeling is a critical bottleneck in the development of computer vision technologies, often constraining machine learning performance due to the time-intensive nature of manual annotations. This work introduces a novel approach that leverages outpainting to mitigate annotated data scarcity by generating artificial contexts and annotations, significantly reducing labeling efforts. We apply this technique to a particularly acute challenge in autonomous driving, urban planning, and environmental monitoring: the lack of diverse, eye-level vehicle images from desired classes. Our dataset comprises AI-generated vehicle images obtained by detecting and cropping vehicles from manually selected seed images, which are then outpainted onto larger canvases to simulate varied real-world conditions. The outpainted images include detailed annotations, providing high-quality ground truth data. Advanced outpainting techniques and image quality assessments ensure visual fidelity and contextual relevance. Ablation results show that incorporating AIDOVECL improves overall detection performance by up to about 10%, and delivers gains of up to about 40% in settings with greater diversity of context, object scale, and placement, with underrepresented classes achieving up to about 50% higher true positives. AIDOVECL enhances vehicle detection by augmenting real training data and supporting evaluation across diverse scenarios. By demonstrating outpainting as an automatic annotation paradigm, it offers a practical and versatile solution for building fine-grained datasets with reduced labeling effort across multiple machine learning domains. The code and links to datasets are available for further research and replication at https://github.com/amir-kazemi/aidovecl.

cs.CV

Uncertainty Quantification in Reduced-Order Gas-Phase Atmospheric Chemistry Modeling using Ensemble SINDy

Uncertainty quantification during atmospheric chemistry modeling is computationally expensive as it typically requires a large number of simulations using complex models. As large-scale modeling is typically performed with simplified chemical mechanisms for computational tractability, we describe a probabilistic surrogate modeling method using principal components analysis (PCA) and Ensemble Sparse Identification of Nonlinear Dynamics (E-SINDy) to both automatically simplify a gas-phase chemistry mechanism and to quantify the uncertainty introduced when doing so. We demonstrate the application of this method on a small photochemical box model for ozone formation. With 100 ensemble members, the calibration $R$-squared value is 0.96 among the three latent species on average and 0.98 for ozone, demonstrating that predicted model uncertainty aligns well with actual model error. In addition to uncertainty quantification, this probabilistic method also improves accuracy as compared to an equivalent deterministic version, by $\sim$60% for the ensemble prediction mean or $\sim$50% for deterministic prediction by the best-performing single ensemble member. Overall, the ozone testing root mean square error (RMSE) is 15.1% of its root mean square (RMS) concentration. Although our probabilistic ensemble simulation ends up being slower than the reference model it emulates, we expect that use of a more complex reference model in future work will result in additional opportunities for acceleration. Versions of this approach applied to full-scale chemical mechanisms may result in improved uncertainty quantification in models of atmospheric composition, leading to enhanced atmospheric understanding and improved support for air quality control and regulation.

physics.comp-ph

Atmospheric chemistry surrogate modeling with sparse identification of nonlinear dynamics

Modeling atmospheric chemistry is computationally expensive and limits the widespread use of atmospheric chemical transport models. This computational cost arises from solving high-dimensional systems of stiff differential equations. Previous work has demonstrated the promise of machine learning (ML) to accelerate air quality model simulations but has suffered from numerical instability during long-term simulations. This may be because previous ML-based efforts have relied on explicit Euler time integration -- which is known to be unstable for stiff systems -- and have used neural networks which are prone to overfitting. We hypothesize that the creation of parsimonious models combined with modern numerical integration techniques can overcome this limitation. Using a small-scale photochemical mechanism to explore the potential of these methods, we have created a machine-learned surrogate by (1) reducing dimensionality using singular value decomposition to create an interpretably-compressed low-dimensional latent space, and (2) using Sparse Identification of Nonlinear Dynamics (SINDy) to create a differential-equation-based representation of the underlying chemical dynamics in the compressed latent space with reduced numerical stiffness. The root mean square error of the ML model prediction for ozone concentration over nine days is 37.8% of the root mean concentration across all simulations in our testing dataset. The surrogate model is 11$\times$ faster with 12$\times$ fewer integration timesteps compared to the reference model and is numerically stable in all tested simulations. Overall, we find that SINDy can be used to create fast, stable, and accurate surrogates of a simple photochemical mechanism. In future work, we will explore the application of this method to more detailed mechanisms and their use in large-scale simulations.

physics.comp-ph

Learned 1-D passive scalar advection to accelerate chemical transport modeling: a case study with GEOS-FP horizontal wind fields

We developed and applied a machine-learned discretization for one-dimensional (1-D) horizontal passive scalar advection, which is an operator component common to all chemical transport models (CTMs). Our learned advection scheme resembles a second-order accuracy, three-stencil numerical solver, but differs from a traditional solver in that coefficients for each equation term are output by a neural network rather than being theoretically-derived constants. We downsampled higher-resolution simulation results -- resulting in up to 16$\times$ larger grid size and 64$\times$ larger timestep -- and trained our neural network-based scheme to match the downsampled integration data. In this way, we created an operator that is low-resolution (in time or space) but can reproduce the behavior of a high-resolution traditional solver. Our model shows high fidelity in reproducing its training dataset (a single 10-day 1-D simulation) and is similarly accurate in simulations with unseen initial conditions, wind fields, and grid spacing. In many cases, our learned solver is more accurate than a low-resolution version of the reference solver, but the low-resolution reference solver achieves greater computational speedup (500$\times$ acceleration) over the high-resolution simulation than the learned solver is able to (18$\times$ acceleration). Surprisingly, our learned 1-D scheme -- when combined with a splitting technique -- can be used to predict 2-D advection, and is in some cases more stable and accurate than the low-resolution reference solver in 2-D. Overall, our results suggest that learned advection operators may offer a higher-accuracy method for accelerating CTM simulations as compared to simply running a traditional integrator at low resolution.

physics.ao-ph

Learned 1-D advection solver to accelerate air quality modeling

Accelerating the numerical integration of partial differential equations by learned surrogate model is a promising area of inquiry in the field of air pollution modeling. Most previous efforts in this field have been made on learned chemical operators though machine-learned fluid dynamics has been a more blooming area in machine learning community. Here we show the first trial on accelerating advection operator in the domain of air quality model using a realistic wind velocity dataset. We designed a convolutional neural network-based solver giving coefficients to integrate the advection equation. We generated a training dataset using a 2nd order Van Leer type scheme with the 10-day east-west components of wind data on 39$^{\circ}$N within North America. The trained model with coarse-graining showed good accuracy overall, but instability occurred in a few cases. Our approach achieved up to 12.5$\times$ acceleration. The learned schemes also showed fair results in generalization tests.

physics.ao-ph

Orders-of-magnitude speedup in atmospheric chemistry modeling through neural network-based emulation

Chemical transport models (CTMs), which simulate air pollution transport, transformation, and removal, are computationally expensive, largely because of the computational intensity of the chemical mechanisms: systems of coupled differential equations representing atmospheric chemistry. Here we investigate the potential for machine learning to reproduce the behavior of a chemical mechanism, yet with reduced computational expense. We create a 17-layer residual multi-target regression neural network to emulate the Carbon Bond Mechanism Z (CBM-Z) gas-phase chemical mechanism. We train the network to match CBM-Z predictions of changes in concentrations of 77 chemical species after one hour, given a range of chemical and meteorological input conditions, which it is able to do with root-mean-square error (RMSE) of less than 1.97 ppb (median RMSE = 0.02 ppb), while achieving a 250x computational speedup. An additional 17x speedup (total 4250x speedup) is achieved by running the neural network on a graphics-processing unit (GPU). The neural network is able to reproduce the emergent behavior of the chemical system over diurnal cycles using Euler integration, but additional work is needed to constrain the propagation of errors as simulation time progresses.

physics.ao-ph