SearcharxivSearch

arXiv subjects

Andreas Marek

Publications and source records attributed to Andreas Marek.

At least 19 recordsLinked to original sources

Beyond Stoner-Wohlfarth: Machine-Learning Models and Symbolic Regression of Hard-Magnet Properties

Predicting the extrinsic properties from hysteresis loops of a magnetic grain, namely the coercive field, remanent magnetisation, and maximum energy product, from its intrinsic micromagnetic parameters is a central problem in permanent-magnet modelling. Established analytical models provide useful estimates but often neglect nonuniform magnetisation processes, whereas direct micromagnetic simulations are computationally expensive. In this work, we train machine-learning models on 12012 micromagnetic simulations of an idealised cubic grain, spanning broad ranges of the saturation magnetisation, exchange constant, and uniaxial anisotropy constant. Benchmarked against the analytical models on identical held-out data, the machine-learning models predict all three extrinsic properties with substantially lower errors. Symbolic regression recovers the Kronm\"uller form of the coercive field, with an effective demagnetising factor that depends on the material, and finds new closed-form expressions for the remanence and maximum energy product. Each law contains at most two fitted constants yet approaches the accuracy of the machine-learning models. We also investigate the inverse problem of recovering the intrinsic parameters from the three extrinsic properties. The saturation magnetisation and anisotropy constant are recovered accurately, whereas the exchange constant is not, because it influences the extrinsic properties only weakly. The trained models are released through the mammos-ai Python package, enabling thousands of candidate parameter sets to be screened in seconds rather than the hours or days required by direct micromagnetic simulation.

cond-mat.str-el

Solvers for Large-Scale Electronic Structure Theory: ELPA and ELSI

In this contribution, we give an overview of the ELPA library and ELSI interface, which are crucial elements for large-scale electronic structure calculations in FHI-aims. ELPA is a key solver library that provides efficient solutions for both standard and generalized eigenproblems, which are central to the Kohn-Sham formalism in density functional theory (DFT). It supports CPU and GPU architectures, with full support for NVIDIA and AMD GPUs, and ongoing development for Intel GPUs. Here we also report the results of recent optimizations, leading to significant improvements in GPU performance for the generalized eigenproblem. ELSI is an open-source software interface layer that creates a well-defined connection between "user" electronic structure codes and "solver" libraries for the Kohn-Sham problem, abstracting the step between Hamilton and overlap matrices (as input to ELSI and the respective solvers) and eigenvalues and eigenvectors or density matrix solutions (as output to be passed back to the "user" electronic structure code). In addition to ELPA, ELSI supports solvers including LAPACK and MAGMA, the PEXSI and NTPoly libraries (which bypass an explicit eigenvalue solution), and several others.

cond-mat.mtrl-sci

3D deep learning for enhanced atom probe tomography analysis of nanoscale microstructures

Quantitative analysis of microstructural features on the nanoscale, including precipitates, local chemical orderings (LCOs) or structural defects (e.g. stacking faults) plays a pivotal role in understanding the mechanical and physical responses of engineering materials. Atom probe tomography (APT), known for its exceptional combination of chemical sensitivity and sub-nanometer resolution, primarily identifies microstructures through compositional segregations. However, this fails when there is no significant segregation, as can be the case for LCOs and stacking faults. Here, we introduce a 3D deep learning approach, AtomNet, designed to process APT point cloud data at the single-atom level for nanoscale microstructure extraction, simultaneously considering compositional and structural information. AtomNet is showcased in segmenting L12-type nanoprecipitates from the matrix in an AlLiMg alloy, irrespective of crystallographic orientations, which outperforms previous methods. AtomNet also allows for 3D imaging of L10-type LCOs in an AuCu alloy, a challenging task for conventional analysis due to their small size and subtle compositional differences. Finally, we demonstrate the use of AtomNet for revealing 2D stacking faults in a Co-based superalloy, without any defected training data, expanding the capabilities of APT for automated exploration of hidden microstructures. AtomNet pushes the boundaries of APT analysis, and holds promise in establishing precise quantitative microstructure-property relationships across a diverse range of metallic materials.

cond-mat.mtrl-sci

Roadmap on Data-Centric Materials Science

Science is and always has been based on data, but the terms "data-centric" and the "4th paradigm of" materials research indicate a radical change in how information is retrieved, handled and research is performed. It signifies a transformative shift towards managing vast data collections, digital repositories, and innovative data analytics methods. The integration of Artificial Intelligence (AI) and its subset Machine Learning (ML), has become pivotal in addressing all these challenges. This Roadmap on Data-Centric Materials Science explores fundamental concepts and methodologies, illustrating diverse applications in electronic-structure theory, soft matter theory, microstructure research, and experimental techniques like photoemission, atom probe tomography, and electron microscopy. While the roadmap delves into specific areas within the broad interdisciplinary field of materials science, the provided examples elucidate key concepts applicable to a wider range of topics. The discussed instances offer insights into addressing the multifaceted challenges encountered in contemporary materials research.

cond-mat.mtrl-sci

Fermionic Quantum Turbulence: Pushing the Limits of High-Performance Computing

Ultracold atoms provide a platform for analog quantum computer capable of simulating the quantum turbulence that underlies puzzling phenomena like pulsar glitches in rapidly spinning neutron stars. Unlike other platforms like liquid helium, ultracold atoms have a viable theoretical framework for dynamics, but simulations push the edge of current classical computers. We present the largest simulations of fermionic quantum turbulence to date and explain the computing technology needed, especially improvements in the ELPA library that enable us to diagonalize matrices of record size (millions by millions). We quantify how dissipation and thermalization proceed in fermionic quantum turbulence by using the internal structure of vortices as a new probe of the local effective temperature. All simulation data and source codes are made available to facilitate rapid scientific progress in the field of ultracold Fermi gases.

cond-mat.quant-gas

Machine learning-enabled tomographic imaging of chemical short-range atomic ordering

In solids, chemical short-range order (CSRO) refers to the self-organisation of atoms of certain species occupying specific crystal sites. CSRO is increasingly being envisaged as a lever to tailor the mechanical and functional properties of materials. Yet quantitative relationships between properties and the morphology, number density, and atomic configurations of CSRO domains remain elusive. Herein, we showcase how machine learning-enhanced atom probe tomography (APT) can mine the near-atomically resolved APT data and jointly exploit the technique's high elemental sensitivity to provide a 3D quantitative analysis of CSRO in a CoCrNi medium-entropy alloy. We reveal multiple CSRO configurations, with their formation supported by state-of-the-art Monte-Carlo simulations. Quantitative analysis of these CSROs allows us to establish relationships between processing parameters and physical properties. The unambiguous characterization of CSRO will help refine strategies for designing advanced materials by manipulating atomic-scale architectures.

cond-mat.mtrl-sci

Convolutional neural network-assisted recognition of nanoscale L12 ordered structures in face-centred cubic alloys

Nanoscale L12-type ordered structures are widely used in face-centred cubic (FCC) alloys to exploit their hardening capacity and thereby improve mechanical properties. These fine-scale particles are typically fully coherent with matrix with the same atomic configuration disregarding chemical species, which makes them challenging to be characterized. Spatial distribution maps (SDMs) are used to probe local order by interrogating the three-dimensional (3D) distribution of atoms within reconstructed atom probe tomography (APT) data. However, it is almost impossible to manually analyse the complete point cloud ($>10$ million) in search for the partial crystallographic information retained within the data. Here, we proposed an intelligent L12-ordered structure recognition method based on convolutional neural networks (CNNs). The SDMs of a simulated L12-ordered structure and the FCC matrix were firstly generated. These simulated images combined with a small amount of experimental data were used to train a CNN-based L12-ordered structure recognition model. Finally, the approach was successfully applied to reveal the 3D distribution of L12-type $δ^\prime$-Al3(LiMg) nanoparticles with an average radius of 2.54 nm in a FCC Al-Li-Mg system. The minimum radius of detectable nanodomain is even down to 5 Å. The proposed CNN-APT method is promising to be extended to recognize other nanoscale ordered structures and even more-challenging short-range ordered phenomena in the near future.

cond-mat.mtrl-sci

GPU-Acceleration of the ELPA2 Distributed Eigensolver for Dense Symmetric and Hermitian Eigenproblems

The solution of eigenproblems is often a key computational bottleneck that limits the tractable system size of numerical algorithms, among them electronic structure theory in chemistry and in condensed matter physics. Large eigenproblems can easily exceed the capacity of a single compute node, thus must be solved on distributed-memory parallel computers. We here present GPU-oriented optimizations of the ELPA two-stage tridiagonalization eigensolver (ELPA2). On top of cuBLAS-based GPU offloading, we add a CUDA kernel to speed up the back-transformation of eigenvectors, which can be the computationally most expensive part of the two-stage tridiagonalization algorithm. We benchmark the performance of this GPU-accelerated eigensolver on two hybrid CPU-GPU architectures, namely a compute cluster based on Intel Xeon Gold CPUs and NVIDIA Volta GPUs, and the Summit supercomputer based on IBM POWER9 CPUs and NVIDIA Volta GPUs. Consistent with previous benchmarks on CPU-only architectures, the GPU-accelerated two-stage solver exhibits a parallel performance superior to the one-stage counterpart. Finally, we demonstrate the performance of the GPU-accelerated eigensolver developed in this work for routine semi-local KS-DFT calculations comprising thousands of atoms.

physics.comp-ph

High Performance Solution of Skew-symmetric Eigenvalue Problems with Applications in Solving the Bethe-Salpeter Eigenvalue Problem

We present a high-performance solver for dense skew-symmetric matrix eigenvalue problems. Our work is motivated by applications in computational quantum physics, where one solution approach to solve the so-called Bethe-Salpeter equation involves the solution of a large, dense, skew-symmetric eigenvalue problem. The computed eigenpairs can be used to compute the optical absorption spectrum of molecules and crystalline systems. One state-of-the art high-performance solver package for symmetric matrices is the ELPA (Eigenvalue SoLvers for Petascale Applications) library. We extend the methods available in ELPA to skew-symmetric matrices. This way, the presented solution method can benefit from the optimizations available in ELPA that make it a well-established, efficient and scalable library, such as GPU support. We compare performance and scalability of our method to the only available high-performance approach for skew-symmetric matrices, an indirect route involving complex arithmetic. In total, we achieve a performance that is up to 3.67 higher than the reference method using Intel's ScaLAPACK implementation. The runtime to solve the Bethe-Salpeter-Eigenvalue problem can be improved by a factor of 10. Our method is freely available in the current release of the ELPA library.

math.NA

Benefits from using mixed precision computations in the ELPA-AEO and ESSEX-II eigensolver projects

We first briefly report on the status and recent achievements of the ELPA-AEO (Eigenvalue Solvers for Petaflop Applications - Algorithmic Extensions and Optimizations) and ESSEX II (Equipping Sparse Solvers for Exascale) projects. In both collaboratory efforts, scientists from the application areas, mathematicians, and computer scientists work together to develop and make available efficient highly parallel methods for the solution of eigenvalue problems. Then we focus on a topic addressed in both projects, the use of mixed precision computations to enhance efficiency. We give a more detailed description of our approaches for benefiting from either lower or higher precision in three selected contexts and of the results thus obtained.

physics.comp-ph

Progenitor-dependent Explosion Dynamics in Self-consistent, Axisymmetric Simulations of Neutrino-driven Core-collapse Supernovae

We present self-consistent, axisymmetric core-collapse supernova simulations performed with the Prometheus-Vertex code for 18 pre-supernova models in the range of 11-28 solar masses, including progenitors recently investigated by other groups. All models develop explosions, but depending on the progenitor structure, they can be divided into two classes. With a steep density decline at the Si/Si-O interface, the arrival of this interface at the shock front leads to a sudden drop of the mass-accretion rate, triggering a rapid approach to explosion. With a more gradually decreasing accretion rate, it takes longer for the neutrino heating to overcome the accretion ram pressure and explosions set in later. Early explosions are facilitated by high mass-accretion rates after bounce and correspondingly high neutrino luminosities combined with a pronounced drop of the accretion rate and ram pressure at the Si/Si-O interface. Because of rapidly shrinking neutron star radii and receding shock fronts after the passage through their maxima, our models exhibit short advection time scales, which favor the efficient growth of the standing accretion-shock instability. The latter plays a supportive role at least for the initiation of the re-expansion of the stalled shock before runaway. Taking into account the effects of turbulent pressure in the gain layer, we derive a generalized condition for the critical neutrino luminosity that captures the explosion behavior of all models very well. We validate the robustness of our findings by testing the influence of stochasticity, numerical resolution, and approximations in some aspects of the microphysics.

astro-ph.SR

Erratum: Progenitor-explosion connection and remnant birth masses for neutrino-driven supernovae of iron-core progenitors (2012, ApJ, 757, 69)

An erroneous interpretation of the hydrodynamical results led to an incorrect determination of the fallback masses in Ugliano et al. (2012), which also (on a smaller level) affects the neutron star masses provided in that paper. This problem was already addressed and corrected in the follow-up works by Ertl et al. (2015) and Sukhbold et al. (2015). Therefore, the reader is advised to use the new data of the latter two publications. In the remaining text of this Erratum we present the differences of the old and new fallback results in detail and explain the origin of the mistake in the original analysis by Ugliano et al. (2012).

astro-ph.HE

Neutrino-driven explosion of a 20 solar-mass star in three dimensions enabled by strange-quark contributions to neutrino-nucleon scattering

Interactions with neutrons and protons play a crucial role for the neutrino opacity of matter in the supernova core. Their current implementation in many simulation codes, however, is rather schematic and ignores not only modifications for the correlated nuclear medium of the nascent neutron star, but also free-space corrections from nucleon recoil, weak magnetism or strange quarks, which can easily add up to changes of several 10% for neutrino energies in the spectral peak. In the Garching supernova simulations with the Prometheus-Vertex code, such sophistications have been included for a long time except for the strange-quark contributions to the nucleon spin, which affect neutral-current neutrino scattering. We demonstrate on the basis of a 20 M_sun progenitor star that a moderate strangeness-dependent contribution of g_a^s = -0.2 to the axial-vector coupling constant g_a = 1.26 can turn an unsuccessful three-dimensional (3D) model into a successful explosion. Such a modification is in the direction of current experimental results and reduces the neutral-current scattering opacity of neutrons, which dominate in the medium around and above the neutrinosphere. This leads to increased luminosities and mean energies of all neutrino species and strengthens the neutrino-energy deposition in the heating layer. Higher nonradial kinetic energy in the gain layer signals enhanced buoyancy activity that enables the onset of the explosion at ~300 ms after bounce, in contrast to the model with vanishing strangeness contributions to neutrino-nucleon scattering. Our results demonstrate the close proximity to explosion of the previously published, unsuccessful 3D models of the Garching group.

astro-ph.SR

Neutrino-driven supernova of a low-mass iron-core progenitor boosted by three-dimensional turbulent convection

We present the first successful simulation of a neutrino-driven supernova explosion in three dimensions (3D), using the Prometheus-Vertex code with an axis-free Yin-Yang grid and a sophisticated treatment of three-flavor, energy-dependent neutrino transport. The progenitor is a nonrotating, zero-metallicity 9.6 Msun star with an iron core. While in spherical symmetry outward shock acceleration sets in later than 300 ms after bounce, a successful explosion starts at ~130 ms postbounce in two dimensions (2D). The 3D model explodes at about the same time but with faster shock expansion than in 2D and a more quickly increasing and roughly 10 percent higher explosion energy of >10^50 erg. The more favorable explosion conditions in 3D are explained by lower temperatures and thus reduced neutrino emission in the cooling layer below the gain radius. This moves the gain radius inward and leads to a bigger mass in the gain layer, whose larger recombination energy boosts the explosion energy in 3D. These differences are caused by less coherent, less massive, and less rapid convective downdrafts associated with postshock convection in 3D. The less violent impact of these accretion downflows in the cooling layer produces less shock heating and therefore diminishes energy losses by neutrino emission. We thus have, for the first time, identified a reduced mass accretion rate, lower infall velocities, and a smaller surface filling factor of convective downdrafts as consequences of 3D postshock turbulence that facilitate neutrino-driven explosions and strengthen them compared to the 2D case.

astro-ph.SR

Self-sustained asymmetry of lepton-number emission: A new phenomenon during the supernova shock-accretion phase in three dimensions

During the stalled-shock phase of our 3D hydrodynamical core-collapse simulations with energy-dependent, 3-flavor neutrino transport, the lepton-number flux (nue minus antinue) emerges predominantly in one hemisphere. This novel, spherical-symmetry breaking neutrino-hydrodynamical instability is termed LESA for "Lepton-number Emission Self-sustained Asymmetry." While the individual nue and antinue fluxes show a pronounced dipole pattern, the heavy-flavor neutrino fluxes and the overall luminosity are almost spherically symmetric. Initially, LESA seems to develop stochastically from convective fluctuations, it exists for hundreds of milliseconds or more, and it persists during violent shock sloshing associated with the standing accretion shock instability. The nue minus antinue flux asymmetry originates mainly below the neutrinosphere in a region of pronounced proto-neutron star (PNS) convection, which is stronger in the hemisphere of enhanced lepton-number flux. On this side of the PNS, the mass-accretion rate of lepton-rich matter is larger, amplifying the lepton-emission asymmetry, because the spherical stellar infall deflects on a dipolar deformation of the stalled shock. The increased shock radius in the hemisphere of less mass accretion and minimal lepton-number flux (antinue flux maximum) is sustained by stronger convection on this side, which is boosted by stronger neutrino heating because the average antinue energy is higher than the average nue energy. Asymmetric heating thus supports the global deformation despite extremely nonstationary convective overturn behind the shock. While these different elements of LESA form a consistent picture, a full understanding remains elusive at present. There may be important implications for neutrino-flavor oscillations, the neutron-to-proton ratio in the neutrino-heated supernova ejecta, and neutron-star kicks, which remain to be explored.

astro-ph.SR

Towards Petaflops Capability of the VERTEX Supernova Code

The VERTEX code is employed for multi-dimensional neutrino-radiation hydrodynamics simulations of core-collapse supernova explosions from first principles. The code is considered state-of-the-art in supernova research and it has been used for modeling for more than a decade, resulting in numerous scientific publications. The computational performance of the code, which is currently deployed on several high-performance computing (HPC) systems up to the Tier-0 class (e.g. in the framework of the European PRACE initiative and the German GAUSS program), however, has so far not been extensively documented. This paper presents a high-level overview of the relevant algorithms and parallelization strategies and outlines the technical challenges and achievements encountered along the evolution of the code from the gigaflops scale with the first, serial simulations in 2000, up to almost petaflops capabilities, as demonstrated lately on the SuperMUC system of the Leibniz Supercomputing Centre (LRZ). In particular, we shall document the parallel scalability and computational efficiency of VERTEX at the large scale and on the major, contemporary HPC platforms. We will outline upcoming scientific requirements and discuss the resulting challenges for the future development and operation of the code.

physics.comp-ph

Porting Large HPC Applications to GPU Clusters: The Codes GENE and VERTEX

We have developed GPU versions for two major high-performance-computing (HPC) applications originating from two different scientific domains. GENE is a plasma microturbulence code which is employed for simulations of nuclear fusion plasmas. VERTEX is a neutrino-radiation hydrodynamics code for "first principles"-simulations of core-collapse supernova explosions. The codes are considered state of the art in their respective scientific domains, both concerning their scientific scope and functionality as well as the achievable compute performance, in particular parallel scalability on all relevant HPC platforms. GENE and VERTEX were ported by us to HPC cluster architectures with two NVidia Kepler GPUs mounted in each node in addition to two Intel Xeon CPUs of the Sandy Bridge family. On such platforms we achieve up to twofold gains in the overall application performance in the sense of a reduction of the time to solution for a given setup with respect to a pure CPU cluster. The paper describes our basic porting strategies and benchmarking methodology, and details the main algorithmic and technical challenges we faced on the new, heterogeneous architecture.

physics.comp-ph

A New Multi-Dimensional General Relativistic Neutrino Hydrodynamics Code of Core-Collapse Supernovae III. Gravitational Wave Signals from Supernova Explosion Models

We present a detailed theoretical analysis of the gravitational-wave (GW) signal of the post-bounce evolution of core-collapse supernovae (SNe), employing for the first time relativistic, two-dimensional (2D) explosion models with multi-group, three-flavor neutrino transport based on the ray-by-ray-plus approximation. The waveforms reflect the accelerated mass motions associated with the characteristic evolutionary stages that were also identified in previous works: A quasi-periodic modulation by prompt postshock convection is followed by a phase of relative quiescence before growing amplitudes signal violent hydrodynamical activity due to convection and the standing accretion shock instability during the accretion period of the stalled shock. Finally, a high-frequency, low-amplitude variation from proto-neutron star (PNS) convection below the neutrinosphere appears superimposed on the low-frequency trend associated with the aspherical expansion of the SN shock after the onset of the explosion. Relativistic effects in combination with detailed neutrino transport are shown to be essential for quantitative predictions of the GW frequency evolution and energy spectrum, because they determine the structure of the PNS surface layer and its characteristic g-mode frequency. Burst-like high-frequency activity phases, correlated with sudden luminosity increase and spectral hardening of electron (anti-)neutrino emission for some 10ms, are discovered as new features after the onset of the explosion. They correspond to intermittent episodes of anisotropic accretion by the PNS in the case of fallback SNe. We find stronger signals for more massive progenitors with large accretion rates. The typical frequencies are higher for massive PNSs, though the time-integrated spectrum also strongly depends on the model dynamics.

astro-ph.SR