SearcharxivSearch

arXiv subjects

Massimo Bernaschi

Publications and source records attributed to Massimo Bernaschi.

At least 19 recordsLinked to original sources

LBFAST: A Lightweight Moment-Represented Lattice Boltzmann Solver for Multi-GPU Architectures

We present LBFAST, a GPU-oriented lattice Boltzmann solver based on a lightweight moment-represented formulation, in which post-collision populations are reconstructed on the fly from a reduced set of moments rather than stored explicitly. This approach significantly lowers the memory footprint, enabling large three-dimensional simulations within the constraints of modern accelerator architectures, where VRAM capacity and bandwidth are critical resources. The method is assessed through standard single- and two-component benchmarks demonstrating good accuracy and stability. Extensive scaling experiments on multi-GPU systems show near-ideal weak scaling up to 512 GPUs and sustained performance across different velocity sets. The combination of reduced memory usage, competitive throughput, and stable energy efficiency makes the proposed formulation a practical route for large-scale lattice Boltzmann simulations on current and emerging HPC platforms.

cs.DC

Scalable s-step Preconditioned Conjugate Gradient with Chebyshev Basis and Gauss-Seidel Gram Solve

We present a variant of the s--step Preconditioned Conjugate Gradient (PCG) method that combines a Chebyshev--stabilized Krylov basis with a Forward Gauss--Seidel (FGS) iteration for the solution of the reduced Gram systems. In s--step PCG, multiple search directions are generated per outer iteration, reducing global synchronization costs but requiring the solution of small dense Gram systems whose conditioning is critical for stability. For a raw Chebyshev Krylov block, we derive an exact moment--based representation of the associated Gram matrix in the unpreconditioned setting and discuss its extension to left preconditioning. This representation provides a structural interpretation of the favorable Gram--matrix behavior observed for moderate step sizes. Building on inexact Krylov theory and on the algebraic connection between FGS and Modified Gram--Schmidt (MGS), we derive residual bounds for the reduced solves and provide a rationale for using a fixed, moderate number of FGS sweeps. Numerical experiments assess the resulting accuracy of the Gram solves and the convergence behavior of the outer iteration. Large--scale experiments on modern NVIDIA GPU architectures, including weak-scaling tests with aggregation-based Algebraic MultiGrid (AMG) preconditioning, show that, within the tested configurations, the proposed Chebyshev--stabilized $s$--step PCG method with FGS Gram solves achieves convergence comparable to classical PCG while reducing global synchronization overhead. The results demonstrate its potential as a scalable communication--reducing solver for current and next--generation accelerator systems.

math.NA

Low energy excitations in a long prism geometry: testing the lower critical dimension of the Ising spin glass

We propose a general method for studying systems that display low-energy excitations in their low-temperature phase. We argue that in a rectangular right-prism geometry (the 3-dimensional generalization of a strip), with one longitudinal size much larger than the transverse size, correlations decay exponentially (at all temperatures) along the longitudinal direction. Still, the scaling of the correlation length with the transverse size carries crucial information from which the lower critical dimension can be inferred. The method is applied in the particularly demanding context of Ising spin glasses at zero magnetic field. The lower critical dimension and the multifractal spectrum for the correlation functions are computed from large-scale numerical simulations. Several technical novelties (such as the unexpectedly crucial performance of Houdayer's cluster method or the convenience of using open---rather than periodic---boundary conditions) allow us to study three-dimensional prisms with transverse dimensions up to $L=24$ and effectively infinite longitudinal dimensions, down to low temperatures. The value that we find for the exponent controlling the behavior of the correlation length agrees with predictions from the Replica Symmetry Breaking (RSB) theory. We argue that our novel setting holds promise in clarifying which of the two competing theories, RSB or the Droplet Model, more accurately describes three-dimensional spin glasses.

cond-mat.dis-nn

Rayleigh-B\'enard thermal convection in emulsions: a short review

Thermally driven emulsions arise in a broad range of natural and industrial contexts, yet their fundamental physical understanding remains only partially established. Emulsions exhibit a complex, concentration-dependent rheology, ranging from Newtonian (dilute emulsions) to yield-stress (concentrated emulsions). In buoyancy-driven flows, the complex structure and rheology of the emulsion are strongly coupled to convective flows, giving rise to fascinating and non-trivial phenomena involving stability, transient dynamics, and morphological evolution of the system. We review recent progress on thermally driven emulsions in the celebrated Rayleigh-B\'enard configuration, offering new perspectives on the behaviour of soft materials in thermal convection.

physics.flu-dyn

On the energy efficiency of sparse matrix computations on multi-GPU clusters

We investigate the energy efficiency of a library designed for parallel computations with sparse matrices. The library leverages high-performance, energy-efficient Graphics Processing Unit (GPU) accelerators to enable large-scale scientific applications. Our primary development objective was to maximize parallel performance and scalability in solving sparse linear systems whose dimensions far exceed the memory capacity of a single node. To this end, we devised methods that expose a high degree of parallelism while optimizing algorithmic implementations for efficient multi-GPU usage. Previous work has already demonstrated the library's performance efficiency on large-scale systems comprising thousands of NVIDIA GPUs, achieving improvements over state-of-the-art solutions. In this paper, we extend those results by providing energy profiles that address the growing sustainability requirements of modern HPC platforms. We present our methodology and tools for accurate runtime energy measurements of the library's core components and discuss the findings. Our results confirm that optimizing GPU computations and minimizing data movement across memory and computing nodes reduces both time-to-solution and energy consumption. Moreover, we show that the library delivers substantial advantages over comparable software frameworks on standard benchmarks.

cs.DC

Role of interfacial stabilization in the Rayleigh-B\'enard convection of liquid-liquid dispersions

Based on mesoscale lattice Boltzmann numerical simulations, we characterize the Rayleigh-B\'enard (RB) convective dynamics of dispersions of liquid droplets in another liquid phase. Our numerical methodology allows us to modify the droplets' interfacial properties to mimic the presence of an emulsifier (e.g., a surfactant), resulting in a positive disjoining pressure that stabilizes the droplets against coalescence. To appreciate the effects of this interfacial stabilization on the RB convective dynamics, we carry out a comparative study between a proper emulsion, i.e., a system where the stabilization mechanism is present (stabilized liquid-liquid dispersion), and a system where the stabilization mechanism is absent (non-stabilized liquid-liquid dispersion). The study is conducted by systematically changing both the volume fraction, $\phi$, and the Rayleigh number, Ra. We find that the morphology of the two systems is dramatically different due to the different interfacial properties. However, the two systems exhibit similar global heat transfer properties, expressed via the Nusselt number Nu. Significant differences in heat transfer emerge at smaller scales, which we analyze via the Nusselt number defined at mesoscales, Nu$_{\mathrm{mes}}$. In particular, stabilized systems exhibit more intense mesoscale heat flux fluctuations due to the persistence of fluid velocity fluctuations down to small scales, which are instead dissipated in the interfacial dynamics of non-stabilized dispersions. For fixed Ra, the difference in mesoscale heat flux fluctuations depends non-trivially on $\phi$, featuring a maximum in the range $0.1 < \phi < 0.2$. Taken all together, our results highlight the role of interfacial physics in mesoscale convective heat transfer of complex fluids.

physics.flu-dyn

Universal exotic dynamics in critical mesoscopic systems: Simulating the square root of Avogadro's number of spins

We explicitly demonstrate the universality of critical dynamics through unprecedented large-scale GPU-based simulations of two out-of-equilibrium processes, comparing the behavior of spin-$1/2$ Ising and spin-$1$ Blume-Capel models on a square lattice. In the first protocol, a completely disordered system is instantaneously brought into contact with a thermal bath at the critical temperature, allowing it to evolve until the coherence length exceeds $10^{3}$ lattice spacings. Finite-size effects are negligible due to the mesoscopic scale of the lattice sizes studied, with linear dimensions up to $L=2^{22}$ and $2^{19}$ for the Ising and Blume-Capel models, respectively. Our numerical data, and the subsequent analysis, demonstrate a strong dynamic universality between the two models and provide the most precise estimate to date of the dynamic critical exponent for this universality class, $z = 2.1676(1)$. In the second protocol, we corroborate the role of the universal ratio of dynamic and static length scales in achieving an exponential acceleration in the approach to equilibrium just \emph{above} the critical temperature, through a time-dependent variation of the thermal bath temperature. The results presented in this work leverage our CUDA-based numerical code, breaking the world record for the simulation speed of the Ising model.

cond-mat.stat-mech

Massive-Scale Simulations of 2D Ising and Blume-Capel Models on Rack-Scale Multi-GPU Systems

We present high-performance implementations of the two-dimensional Ising and Blume-Capel models for large-scale, multi-GPU simulations. Our approach takes full advantage of the NVIDIA GB200 NVL72 system, which features up to $72$ GPUs interconnected via high-bandwidth NVLink, enabling direct GPU-to-GPU memory access across multiple nodes. By utilizing Fabric Memory and an optimized Monte Carlo kernel for the Ising model, our implementation supports simulations of systems with linear sizes up to $L=2^{23}$, corresponding to approximately $70$ trillion spins. This allows for a peak processing rate of nearly $1.15 \times 10^5$ lattice updates per nanosecond-setting a new performance benchmark for Ising model simulations. Additionally, we introduce a custom protocol for computing correlation functions, which strikes an optimal balance between computational efficiency and statistical accuracy. This protocol enables large-scale simulations without incurring prohibitive runtime costs. Benchmark results show near-perfect strong and weak scaling up to $64$ GPUs, demonstrating the effectiveness of our approach for large-scale statistical physics simulations.

physics.comp-ph

Communication-reduced Conjugate Gradient Variants for GPU-accelerated Clusters

Linear solvers are key components in any software platform for scientific and engineering computing. The solution of large and sparse linear systems lies at the core of physics-driven numerical simulations relying on partial differential equations (PDEs) and often represents a significant bottleneck in datadriven procedures, such as scientific machine learning. In this paper, we present an efficient implementation of the preconditioned s-step Conjugate Gradient (CG) method, originally proposed by Chronopoulos and Gear in 1989, for large clusters of Nvidia GPU-accelerated computing nodes. The method, often referred to as communication-reduced or communication-avoiding CG, reduces global synchronizations and data communication steps compared to the standard approach, enhancing strong and weak scalability on parallel computers. Our main contribution is the design of a parallel solver that fully exploits the aggregation of low-granularity operations inherent to the s-step CG method to leverage the high throughput of GPU accelerators. Additionally, it applies overlap between data communication and computation in the multi-GPU sparse matrix-vector product. Experiments on classic benchmark datasets, derived from the discretization of the Poisson PDE, demonstrate the potential of the method.

math.NA

Dynamical regimes of thermally convective emulsions

Emulsions are paramount in various interdisciplinary topical areas, yet a satisfactory understanding of their behavior in buoyancy-driven thermal flows has not been established. In the present work, we unravel the dynamical regimes of thermal convection in emulsions by leveraging a large set of mesoscale numerical simulations. Emulsions are prepared with a given volume fraction of the initially dispersed phase, $\phi$, ranging from dilute (low values of $\phi$) to jammed emulsions (high values of $\phi$), resulting in different rheological responses of the emulsion, i.e., from Newtonian to non-Newtonian yield-stress behaviors, respectively. We then characterize the dynamics of the emulsions in the paradigmatic setup of the Rayleigh-B\'enard convection, i.e., when confined between two parallel walls at different temperatures under the effect of buoyancy forces, the latter encoded in the dimensionless Rayleigh number Ra. We thoroughly investigated the dynamics of the emulsion in the changing of $\phi$ and Ra. For a given $\phi$, at increasing Ra, we observe that the emulsion exhibits convection states, where structural changes may appear (i.e., droplet breakup, coalescence or phase inversion), which inevitably impact the emulsion rheology. For sufficiently high values of Ra, two states of convection are observed: for low/moderate values of $\phi$ (Newtonian emulsions), we observe breakup-dominated dynamics, whereas for high values of $\phi$ (non-Newtonian emulsions), we observe phase-inverted states. For both scenarios, the droplet size distribution depends on Ra, and scaling laws for the average droplet size are analyzed and quantified. Our results offer insights into the rich dynamics of emulsions under thermal convection, offering the first detailed characterization of the various dynamic regimes to be expected and their relation with structural changes occurring in such complex fluids.

physics.flu-dyn

Intermittent thermal convection in jammed emulsions

We study the process of thermal convection in jammed emulsions with a yield-stress rheology. We find that heat transfer occurs via an intermittent mechanism, whereby intense short-lived convective "heat bursts" are spaced out by long-lasting conductive periods. This behaviour is the result of a sequence of fluidization-rigidity transitions, rooted in a non-trivial interplay between emulsion yield-stress rheology and plastic activity, which we characterize via a statistical analysis of the dynamics at the droplet scale. We also show that droplets' coalescence induced during heat bursts leads to a spatially heterogeneous phase-inversion of the emulsion which eventually supports a sustained convective state.

physics.flu-dyn

The QISG suite: high-performance codes for studying Quantum Ising Spin Glasses

We release a set of GPU programs for the study of the Quantum ($S=1/2$) Spin Glass on a square lattice, with binary couplings. The library contains two main codes: MCQSG (that carries out Monte Carlo simulations using both the Metropolis and the Parallel Tempering algorithms, for the problem formulated in the Trotter-Suzuki approximation), and EDQSG (that obtains the extremal eigenvalues of the Transfer Matrix using the Lanczos algorithm). EDQSG has allowed us to diagonalize transfer matrices with size up to $2^{36}\times2^{36}$. From its side, MCQSG running on four NVIDIA A100 cards delivers a sub-picosecond time per spin-update, a performance that is competitive with dedicated hardware. We include as well in our library GPU programs for the analysis of the spin configurations generated by MCQSG. Finally, we provide two auxiliary codes: the first generates the lookup tables employed by the random number generator of MCQSG; the second one simplifies the execution of multiple runs using different input data.

physics.comp-ph

The Quantum Transition of the Two-Dimensional Ising Spin Glass: A Tale of Two Gaps

Quantum annealers are commercial devices aiming to solve very hard computational problems named spin glasses. Just like in metallurgic annealing one slowly cools a ferrous metal, quantum annealers seek good solutions by slowly removing the transverse magnetic field at the lowest possible temperature. The field removal diminishes quantum fluctuations but forces the system to traverse the critical point that separates the disordered phase (at large fields) from the spin-glass phase (at small fields). A full understanding of this phase transition is still missing. A debated, crucial question regards the closing of the energy gap separating the ground state from the first excited state. All hopes of achieving an exponential speed-up, as compared to classical computers, rest on the assumption that the gap will close algebraically with the number of qspins, but renormalization group calculations predict that the closing will be instead exponential. Here we solve this debate through extreme-scale numerical simulations, finding that both parties grasped parts of the truth. While the closing of the gap at the critical point is indeed super-algebraic, it remains algebraic if one restricts the symmetry of possible excitations. Since this symmetry restriction is experimentally achievable (at least nominally), there is still hope for the Quantum Annealing paradigm.

cond-mat.dis-nn

Analysis of the heat transfer fluctuations in the Rayleigh-B\'enard convection of concentrated emulsions with finite-size droplets

Employing numerical simulations, we provide an accurate insight into the of heat transfer mechanisms in the Rayleigh-B\'enard convection of concentrated emulsions with finite-size droplets. We focus on the unsteady dynamics characterizing the thermal convection of these complex fluids close to the transition from conductive to convective states, where the heat transfer phenomenon, expressed in terms of the Nusselt number Nu, is characterized by pronounced fluctuations triggered by collective droplets motion [Pelusi et al., Soft Matter 17(13), 3709 - 3721 (2021)]. By systematically increasing the droplet concentration, we show how these fluctuations emerge along with the segregation of "extreme events" in the boundary layers, causing intermittent bursts in the heat flux fluctuations. Furthermore, we quantify the extension $S$ and the duration $\mathcal{T}$ of the coherent droplet motion accompanying these extreme events via a suitable statistical analysis involving the droplets displacements. We show how the increase in droplet concentration results in a power-law behaviour of the probability distribution function of $S$ and $\mathcal{T}$ and how this outcome is robust at changing the analysis protocol. Our work offers a comprehensive picture, linking macroscopic heat transfer fluctuations with the statistics of droplets at the mesoscale.

physics.flu-dyn

Multi-GPU aggregation-based AMG preconditioner for iterative linear solvers

We present and release in open source format a sparse linear solver which efficiently exploits heterogeneous parallel computers. The solver can be easily integrated into scientific applications that need to solve large and sparse linear systems on modern parallel computers made of hybrid nodes hosting NVIDIA Graphics Processing Unit (GPU) accelerators. The work extends our previous efforts in the exploitation of a single GPU accelerator and proposes an implementation, based on the hybrid MPI-CUDA software environment, of a Krylov-type linear solver relying on an efficient Algebraic MultiGrid (AMG) preconditioner already available in the BootCMatchG library. Our design for the hybrid implementation has been driven by the best practices for minimizing data communication overhead when multiple GPUs are employed, yet preserving the efficiency of the single GPU kernels. Strong and weak scalability results on well-known benchmark test cases of the new version of the library are discussed. Comparisons with the Nvidia AmgX solution show an improvement of up to 2.0x in the solve phase.

cs.DC

Lightweight Lattice Boltzmann

A GPU-accelerated version of the lattice Boltzmann method for efficient simulation of soft materials is introduced. Unlike standard approaches, this method reconstructs the distribution functions from available hydrodynamic variables (density, momentum, and pressure tensor) without storing the full set of discrete populations. This scheme shows satisfactory numerical stability, significantly lower memory requirements, and data access cost. A series of benchmark tests of relevance to soft matter, such as collisions of fluid droplets, is discussed to validate the method. The results can be of particular interest for high-performance simulations of soft matter systems on future exascale computers.

physics.flu-dyn

The Fitness-Corrected Block Model, or how to create maximum-entropy data-driven spatial social networks

Models of networks play a major role in explaining and reproducing empirically observed patterns. Suitable models can be used to randomize an observed network while preserving some of its features, or to generate synthetic graphs whose properties may be tuned upon the characteristics of a given population. In the present paper, we introduce the Fitness-Corrected Block Model, an adjustable-density variation of the well-known Degree-Corrected Block Model, and we show that the proposed construction yields a maximum entropy model. When the network is sparse, we derive an analytical expression for the degree distribution of the model that depends on just the constraints and the chosen fitness-distribution. Our model is perfectly suited to define maximum-entropy data-driven spatial social networks, where each block identifies vertices having similar position (e.g., residence) and age, and where the expected block-to-block adjacency matrix can be inferred from the available data. In this case, the sparse-regime approximation coincides with a phenomenological model where the probability of a link binding two individuals is directly proportional to their sociability and to the typical cohesion of their age-groups, whereas it decays as an inverse-power of their geographic distance. We support our analytical findings through simulations of a stylized urban area.

physics.soc-ph

Blocking Techniques for Sparse Matrix Multiplication on Tensor Accelerators

Tensor accelerators have gained popularity because they provide a cheap and efficient solution for speeding up computational-expensive tasks in Deep Learning and, more recently, in other Scientific Computing applications. However, since their features are specifically designed for tensor algebra (typically dense matrix-product), it is commonly assumed that they are not suitable for applications with sparse data. To challenge this viewpoint, we discuss methods and present solutions for accelerating sparse matrix multiplication on such architectures. In particular, we present a 1-dimensional blocking algorithm with theoretical guarantees on the density, which builds dense blocks from arbitrary sparse matrices. Experimental results show that, even for unstructured and highly-sparse matrices, our block-based solution which exploits Nvidia Tensor Cores is faster than its sparse counterpart. We observed significant speed-ups of up to two orders of magnitude on real-world sparse matrices.

cs.DC