SearcharxivSearch

arXiv subjects

Carlo Graziani

Publications and source records attributed to Carlo Graziani.

At least 19 recordsLinked to original sources

Application Failures and Machine Computational Efficiency

We present a framework for evaluating uptime efficiency of Exascale-class scientific computers when application failure rates are appreciable. This is the situation that confronts current leadership-class scientific computing platforms and large AI training installations. What distinguishes scientific computing platforms is the heterogeneity of their applications. We argue that this diversity requires that failure rates and mean intervals between failures should be specified in terms of \emph{usage} (e.g. node-hours) rather than time, as is currently customary. We consider the usage loss terms due to failures, to checkpointing, and to restart costs, and update the framework of Daly (2006) allowing users to specify optimal checkpointing usage intervals that minimize such losses. We derive the machine computational efficiency, which specifies the expected fractional resource allocation that is available for scientific computation. We illustrate the methodology using one year of production runtime data from the \emph{Frontier} supercomputer at Oak Ridge National Laboratory.

cs.DC

A Scalable Pattern Mining Workflow for Interpretable Machine Log Analysis in High-Performance Computing Environments

Modern supercomputers housed in High Performance Computing (HPC) environments generate massive volumes of log data daily, revealing intricate information and performance metrics about these complex systems. The sheer size and heterogeneous nature of HPC logs, especially text data, pose significant challenges for traditional analytical techniques. Consequently, more complex workflows are necessary for pattern extraction when analyzing these logs, enabling the discovery of underlying patterns and anomalies that may indicate system faults and help predict future failures and inefficiencies. Our log analysis workflow investigates a combination of advanced pattern-matching and mining techniques applied to HPC log analysis. By systematically identifying frequent log patterns and pattern sequences in log messages and storing them in a finite-state automaton, such as the Aho-Corasick automaton, our workflow enables automated detection of frequent errors and fault events. To extract these patterns and sequences, we leverage information about system hierarchy and message priority. We then correlate and cluster the identified error sequences with job logs, revealing groups of applications with similar or dissimilar error signatures. This approach yields insights that inform improvements and guide real-time monitoring efforts. Our research establishes that pattern mining is vital for unlocking the full potential of log data by enabling real-time analysis and contributing to more resilient, scalable HPC systems. We demonstrate the effectiveness of our approach through summary statistics and a case study on an exascale-class system supercomputer.

cs.DC

Tomography of Atomic Nuclei

We carry out continuum quantum Monte Carlo calculations of the quantum-mechanical Wigner distribution functions of selected nuclei, up to $^{16}$O. These distributions provide a form of quantum tomography of the spatial and momentum structure of the system. They also help identify the location of high-momentum regions in atomic nuclei and provide insight into the onset of alpha clustering. Besides their intrinsic interest, these distributions will be useful for neutrino event generators, as they correlate the positions and momenta of nucleons in the initial target state. To facilitate their application, we address the need to store them compactly by developing an accurate Gaussian Process emulator that automatically preserves their normalization.

nucl-th

Probabilistic Attribution For Large Language Models

The generative nature of Large Language Models (LLMs) is reflected in the conditional probabilities they compute to sample each response token given the previous tokens. These probabilities encode the distributional structure that the model learns in training and exploits in inference. In this work, we use these probabilities to situate LLMs within the mathematical theory of stochastic processes. We use this framework to design a model-agnostic probabilistic token attribution measure, using Bayes rule to invert the next-token log-probabilities so as to capture the models internal representation of the distribution over token sequences. The representation is independent of the models computational structure. This representation yields the conditional probability of the response given the prompt, and of the response given the prompt with a token marginalized away. Our attribution score is the log of the ratio of these probabilities. We further compute the entropies of a single prompts token distributions, conditioned on the remaining context. The interplay between entropy and attribution score sheds light on LLM behavior. We evaluate 8 models across 7 prompts and investigate anomalies, token sensitivity, response stability, model stability, and training convergence, thereby improving interpretability and guiding users to focus on uncertain or unstable parts of the generation.

cs.CL

Kernel Model Validation: How To Do It, And Why You Should Care

Gaussian Process (GP) models are popular tools in uncertainty quantification (UQ) because they purport to furnish functional uncertainty estimates that can be used to represent model uncertainty. It is often difficult to state with precision what probabilistic interpretation attaches to such an uncertainty, and in what way is it calibrated. Without such a calibration statement, the value of such uncertainty estimates is quite limited and qualitative. We motivate the importance of proper probabilistic calibration of GP predictions by describing how GP predictive calibration failures can cause degraded convergence properties in a target optimization algorithm called Targeted Adaptive Design (TAD). We discuss the interpretation of GP-generated uncertainty intervals in UQ, and how one may learn to trust them, through a formal procedure for covariance kernel validation that exploits the multivariate normal nature of GP predictions. We give simple examples of GP regression misspecified 1-dimensional models, and discuss the situation with respect to higher-dimensional models.

stat.ME

Targeted Adaptive Design

Modern advanced manufacturing and advanced materials design often require searches of relatively high-dimensional process control parameter spaces for settings that result in optimal structure, property, and performance parameters. The mapping from the former to the latter must be determined from noisy experiments or from expensive simulations. We abstract this problem to a mathematical framework in which an unknown function from a control space to a design space must be ascertained by means of expensive noisy measurements, which locate optimal control settings generating desired design features within specified tolerances, with quantified uncertainty. We describe targeted adaptive design (TAD), a new algorithm that performs this sampling task efficiently. TAD creates a Gaussian process surrogate model of the unknown mapping at each iterative stage, proposing a new batch of control settings to sample experimentally and optimizing the updated log-predictive likelihood of the target design. TAD either stops upon locating a solution with uncertainties that fit inside the tolerance box or uses a measure of expected future information to determine that the search space has been exhausted with no solution. TAD thus embodies the exploration-exploitation tension in a manner that recalls, but is essentially different from, Bayesian optimization and optimal experimental design.

cs.LG

OS-net: Orbitally Stable Neural Networks

We introduce OS-net (Orbitally Stable neural NETworks), a new family of neural network architectures specifically designed for periodic dynamical data. OS-net is a special case of Neural Ordinary Differential Equations (NODEs) and takes full advantage of the adjoint method based backpropagation method. Utilizing ODE theory, we derive conditions on the network weights to ensure stability of the resulting dynamics. We demonstrate the efficacy of our approach by applying OS-net to discover the dynamics underlying the Rössler and Sprott's systems, two dynamical systems known for their period doubling attractors and chaotic behavior.

math.DS

Simultaneous Reconstruction and Uncertainty Quantification for Tomography

Tomographic reconstruction, despite its revolutionary impact on a wide range of applications, suffers from its ill-posed nature in that there is no unique solution because of limited and noisy measurements. Therefore, in the absence of ground truth, quantifying the solution quality is highly desirable but under-explored. In this work, we address this challenge through Gaussian process modeling to flexibly and explicitly incorporate prior knowledge of sample features and experimental noises through the choices of the kernels and noise models. Our proposed method yields not only comparable reconstruction to existing practical reconstruction methods (e.g., regularized iterative solver for inverse problem) but also an efficient way of quantifying solution uncertainties. We demonstrate the capabilities of the proposed approach on various images and show its unique capability of uncertainty quantification in the presence of various noises.

stat.AP

A Deep Learning Approach to Probabilistic Forecasting of Weather

We discuss an approach to probabilistic forecasting based on two chained machine-learning steps: a dimensional reduction step that learns a reduction map of predictor information to a low-dimensional space in a manner designed to preserve information about forecast quantities; and a density estimation step that uses the probabilistic machine learning technique of normalizing flows to compute the joint probability density of reduced predictors and forecast quantities. This joint density is then renormalized to produce the conditional forecast distribution. In this method, probabilistic calibration testing plays the role of a regularization procedure, preventing overfitting in the second step, while effective dimensional reduction from the first step is the source of forecast sharpness. We verify the method using a 22-year 1-hour cadence time series of Weather Research and Forecasting (WRF) simulation data of surface wind on a grid.

stat.ML

An Application of Gaussian Process Modeling for High-order Accurate Adaptive Mesh Refinement Prolongation

We present a new polynomial-free prolongation scheme for Adaptive Mesh Refinement (AMR) simulations of compressible and incompressible computational fluid dynamics. The new method is constructed using a multi-dimensional kernel-based Gaussian Process (GP) prolongation model. The formulation for this scheme was inspired by the GP methods introduced by A. Reyes et al. (A New Class of High-Order Methods for Fluid Dynamics Simulation using Gaussian Process Modeling, Journal of Scientific Computing, 76 (2017), 443-480; A variable high-order shock-capturing finite difference method with GP-WENO, Journal of Computational Physics, 381 (2019), 189-217). In this paper, we extend the previous GP interpolations and reconstructions to a new GP-based AMR prolongation method that delivers a high-order accurate prolongation of data from coarse to fine grids on AMR grid hierarchies. In compressible flow simulations special care is necessary to handle shocks and discontinuities in a stable manner. To meet this, we utilize the shock handling strategy using the GP-based smoothness indicators developed in the previous GP work by A. Reyes et al. We demonstrate the efficacy of the GP-AMR method in a series of testsuite problems using the AMReX library, in which the GP-AMR method has been implemented.

math.NA

Probabilistic Recalibration of Forecasts

We present a scheme by which a probabilistic forecasting system whose predictions have poor probabilistic calibration may be recalibrated by incorporating past performance information to produce a new forecasting system that is demonstrably superior to the original, in that one may use it to consistently win wagers against someone using the original system. The scheme utilizes Gaussian process (GP) modeling to estimate a probability distribution over the Probability Integral Transform (PIT) of a scalar predictand. The GP density estimate gives closed-form access to information entropy measures associated with the estimated distribution, which allows prediction of winnings in wagers against the base forecasting system. A separate consequence of the procedure is that the recalibrated forecast has a uniform expected PIT distribution. A distinguishing feature of the procedure is that it is appropriate even if the PIT values are not i.i.d. The recalibration scheme is formulated in a framework that exploits the deep connections between information theory, forecasting, and betting. We demonstrate the effectiveness of the scheme in two case studies: a laboratory experiment with a nonlinear circuit and seasonal forecasts of the intensity of the El Niño-Southern Oscillation phenomenon.

stat.ME

A variable high-order shock-capturing finite difference method with GP-WENO

We present a new finite difference shock-capturing scheme for hyperbolic equations on static uniform grids. The method provides selectable high-order accuracy by employing a kernel-based Gaussian Process (GP) data prediction method which is an extension of the GP high-order method originally introduced in a finite volume framework by the same authors. The method interpolates Riemann states to high order, replacing the conventional polynomial interpolations with polynomial-free GP-based interpolations. For shocks and discontinuities, this GP interpolation scheme uses a nonlinear shock handling strategy similar to Weighted Essentially Non-oscillatory (WENO), with a novelty consisting in the fact that nonlinear smoothness indicators are formulated in terms of the Gaussian likelihood of the local stencil data, replacing the conventional $L_2$-type smoothness indicators of the original WENO method. We demonstrate that these GP-based smoothness indicators play a key role in the new algorithm, providing significant improvements in delivering high -- and selectable -- order accuracy in smooth flows, while successfully delivering non-oscillatory solution behavior in discontinuous flows.

physics.comp-ph

A New Class of High-Order Methods for Fluid Dynamics Simulations using Gaussian Process Modeling

We introduce an entirely new class of high-order methods for computational fluid dynamics (CFD) based on the Gaussian Process (GP) family of stochastic functions. Our approach is to use kernel-based GP prediction methods to interpolate/reconstruct high-order approximations for solving hyperbolic PDEs. We present the GP approach as a new formulation of high-order (magneto)hydrodynamic state variable interpolation that furnishes an alternative to conventional polynomial-based approaches.

physics.comp-ph

Inferring Morphology and Strength of Magnetic Fields From Proton Radiographs

Proton radiography is an important diagnostic method for laser plasma experiments, and is particularly important in the analysis of magnetized plasmas. The theory of radiographic image analysis has heretofore only permitted somewhat limited analysis of the radiographs of such plasmas. We furnish here a theory that remedies this deficiency. We show that to linear order in magnetic field gradients, proton radiographs are projection images of the MHD current along the proton trajectories. We demonstrate that in the linear approximation, the full structure of the perpedicular magnetic field can be reconstructed by solving a steady-state inhomogeneous 2-dimensional diffusion equation sourced by the radiograph fluence contrast data. We explore limitations of the inversion method due to Poisson noise, to discretization errors, to radiograph edge effects, and to obstruction by laser target structures. We also provide a separate analysis that is well-suited to the inference of isotropic-homogeneous magnetic turbulence spectra. We discuss extension of these results to the nonlinear contrast regime.

physics.plasm-ph

Probing the Cosmic Gamma-Ray Burst Rate with Trigger Simulations of the Swift Burst Alert Telescope

The gamma-ray burst (GRB) rate is essential for revealing the connection between GRBs, supernovae and stellar evolution. Additionally, the GRB rate at high redshift provides a strong probe of star formation history in the early universe. While hundreds of GRBs are observed by Swift, it remains difficult to determine the intrinsic GRB rate due to the complex trigger algorithm of Swift. Current studies of the GRB rate usually approximate the Swift trigger algorithm by a single detection threshold. However, unlike the previously flown GRB instruments, Swift has over 500 trigger criteria based on photon count rate and additional image threshold for localization. To investigate possible systematic biases and explore the intrinsic GRB properties, we develop a program that is capable of simulating all the rate trigger criteria and mimicking the image threshold. Our simulations show that adopting the complex trigger algorithm of Swift increases the detection rate of dim bursts. As a result, our simulations suggest bursts need to be dimmer than previously expected to avoid over-producing the number of detections and to match with Swift observations. Moreover, our results indicate that these dim bursts are more likely to be high redshift events than low-luminosity GRBs. This would imply an even higher cosmic GRB rate at large redshifts than previous expectations based on star-formation rate measurements, unless other factors, such as the luminosity evolution, are taken into account. The GRB rate from our best result gives a total number of 4571^{+829}_{-1584} GRBs per year that are beamed toward us in the whole universe. SPECIAL NOTE (2015.05.16): This new version incorporates an erratum. All the GRB rate normalizations ($R_{\rm GRB}(z=0)$) should be a factor of 2 smaller than previously reported. Please refer to the Appendix for more details. We sincerely apologize for the mistake.

astro-ph.HE

The Biermann Catastrophe in Numerical MHD

The Biermann Battery effect is frequently invoked in cosmic magnetogenesis and studied in High-Energy Density laboratory physics experiments. Generation of magnetic fields by the Biermann effect due to mis-aligned density and temperature gradients in smooth flow behind shocks is well known. We show that a Biermann-effect magnetic field is also generated within shocks. Direct implementation of the Biermann effect in MHD codes does not capture this physical process, and worse, produces unphysical magnetic fields at shocks whose value does not converge with resolution. We show that this convergence breakdown is due to naive discretization, which fails to account for the fact that discretized irrotational vector fields have spurious solenoidal components that grow without bound near a discontinuity. We show that careful consideration of the kinetics of ion viscous shocks leads to a formulation of the Biermann effect that gives rise to a convergent algorithm. We note two novel physical effects: a resistive magnetic precursor in which Biermann-generated field in the shock "leaks" resistively upstream; and a thermal magnetic precursor , in which field is generated by the Biermann effect ahead of the shock front due to gradients created by the shock's electron thermal conduction precursor. Both effects appear to be potentially observable in experiments at laser facilities. We re-examine published studies of magnetogenesis in galaxy cluster formation, and conclude that the simulations in question had inadequate resolution to reliably estimate the field generation rate. Corrected estimates suggest primordial field values in the range $B\sim 10^{-22}$G --- $10^{-19}$G by $z=3$.

astro-ph.GA

Characterizing the Velocity Fields in Massive Stars

We apply the mathematical formalism of vector spherical harmonics decomposition to convective stellar velocity fields from multi-dimensional hydrodynamics simulations, and show that the resulting power spectra furnish a robust and stable statistical description of stellar convective turbulence. Analysis of the power spectra help identify key physical parameters of the convective process such as the dominant scale of the turbulent motions that influence the structure of massive evolved pre-supernova stars. We introduce the numerical method that can be used to calculate vector spherical harmonics power spectra from 2D and 3D convective shell simulation data. Using this method we study the properties of oxygen shell burning and convection for a 15 Msun star simulated by the hydrodynamics code FLASH in 2D and 3D. We discuss the importance of realistic initial conditions to achieving successful core-collapse supernova explosions in multi-dimensional simulations. We show that the calculated power spectra can be used to generate realizations of the velocity fields of pre-supernova convective shells. We find that the slope of the solenoidal mode power spectrum remains mostly constant throughout the evolution of convection in the oxygen shell in both 2D and 3D simulations. We also find that the characteristic radial scales of the convective elements are smaller in 3D than in 2D while the angular scales are larger in 3D.

astro-ph.SR

Three-dimensional Simulations of Pure Deflagration Models for Thermonuclear Supernovae

We present a systematic study of the pure deflagration model of Type Ia supernovae using three-dimensional, high-resolution, full-star hydrodynamical simulations, nucleosynthetic yields calculated using Lagrangian tracer particles, and light curves calculated using radiation transport. We evaluate the simulations by comparing their predicted light curves with many observed SNe Ia using the SALT2 data-driven model and find that the simulations may correspond to under-luminous SNe Iax. We explore the effects of the initial conditions on our results by varying the number of randomly selected ignition points from 63 to 3500, and the radius of the centered sphere they are confined in from 128 to 384 km. We find that the rate of nuclear burning depends on the number of ignition points at early times, the density of ignition points at intermediate times, and the radius of the confining sphere at late times. The results depend primarily on the number of ignition points, but we do not expect this to be the case in general. The simulations with few ignition points release more nuclear energy $E_{\mathrm{nuc}}$, have larger kinetic energies $E_{\mathrm{K}}$, and produce more $^{56}$Ni than those with many ignition points, and differ in the distribution of $^{56}$Ni, Si, and C/O in the ejecta. For these reasons, the simulations with few ignition points exhibit higher peak B-band absolute magnitudes $M_\mathrm{B}$ and light curves that rise and decline more quickly; their $M_\mathrm{B}$ and light curves resemble those of under-luminous SNe Iax, while those for simulations with many ignition points are not.

astro-ph.HE