SearcharxivSearch

arXiv subjects

Klaus Steiniger

Publications and source records attributed to Klaus Steiniger.

13 recordsLinked to original sources

Monte-Carlo Event Generation for X-Ray Thomson Scattering Analysis

A key diagnostic in warm-dense matter (WDM) experiments is X-ray Thomson scattering (XRTS), but its interpretation is often limited by complex instrument effects and the high computationally expensive combinations of microscopic models with detector simulations. We present a proof-of-principle implementation of an event-driven approach to XRTS modelling, inspired by particle physics event-generators. Instead of computing the spectra via forward models, individual scattering events are sampled from the differential cross section and sent through a spectrometer simulation. This provides a statistically consistent representation that preserves full kinematic information and enables flexible and geometry-aware analysis. We demonstrate the feasibility and physical consistency of the method for non-resonant XRTS in a synthetic setup. By decoupling event generation from detector-level analysis, the framework allows efficient reuse of the sampled events and reduces computational overhead associated with repeated evaluations. The method is model-agnostic and establishes a new connection between particle-physics event generation techniques and WDM diagnostics, providing a scalable foundation for advanced XRTS analysis and inference.

hep-ph

PIConGPU modeling of nanoplasma formation in helium nanodroplets irradiated by intense femtosecond laser pulses

Helium nanodroplets provide a unique and versatile platform for investigating strong-field-driven nanoplasma dynamics. In this work, we present large-scale, GPU-accelerated particle-in-cell simulations using \textsc{PIConGPU} to study the interaction of pure helium nanodroplets containing up to $10^{6}$ atoms with intense near-infrared femtosecond laser pulses, and compare the results with single-shot velocity-map electron imaging and ion measurements. The simulations describe the plasma evolution from the first ionization events to collective electron motion, nanoplasma formation, and early expansion. We show that the calculated electron and ion observables reproduce the main features of the measured spectra in systems with similar cluster sizes and laser intensities. Our results demonstrate that \textsc{PIConGPU} captures the essential physics of nanoplasma formation previously addressed mainly with molecular-dynamics or TDDFT approaches, while remaining computationally efficient and applicable to much larger systems. This establishes \textsc{PIConGPU} as a powerful and scalable tool for connecting nanoplasma theory with experimentally accessible observables.

physics.plasm-ph

Kinetic Simulations of Laser-Driven Compression and Heating of Magnetised Cryogenic Hydrogen Targets using PIConGPU

We present fully kinetic two-dimensional, three-velocity-component (2D3V) PIConGPU simulations of a three-beam direct-drive interaction with a 15 $μ$m solid-density cryogenic hydrogen cylinder, establishing a predictive numerical baseline for the operational DRACO ($τ=30$ fs) and upcoming PENELOPE ($τ=150$ fs) laser facilities at HZDR. The simulations resolve charge-separation fields on the order of 3 TV/m and reveal a robust kinematic bifurcation of the accelerated population into a fast (1-5 MeV) ion beam and a slower bulk (1-100 keV) flow. We demonstrate analytically and numerically that the charge-separation front ($v_{hb}$) is an intrinsically non-quasi-neutral electrostatic double layer that lies outside the closure assumptions of radiation-hydrodynamic models. A simple $2v_{hb}$ reflection scaling derived directly from the front trajectory tracks the centroid of the constant-energy fast-ion band under the impulsive 30 fs driver and the time-varying upper edge of the swept fast-ion band under the sustained 150 fs driver, across both intensities ($a_{0}=12.7$ and 22.0), establishing this non-thermal mechanism as the dominant acceleration pathway. We then scan an external axial magnetic field from 0 T to 10 kT. Laboratory-achievable 20 T fields leave all macroscopic observables unchanged; fields at the kT scale progressively magnetise the MeV hot-electron population, quench the laser-driven charge-separation mechanism, suppress the fast-ion band, and more than double the net-inward compression time of the short-pulse driver-while extending the outer target envelope. A geometric equivalence argument maps these kT-scale results onto larger-diameter cryogenic hydrogen jets.

physics.plasm-ph

Reduction of SAXS Signal due to Doppler Broadening Induced Loss of Coherence

We present an analytical and numerical study of how Doppler-induced spectral broadening in laser-heated plasmas degrades the coherence of small-angle X-ray scattering (SAXS) signals, and show that the resulting loss of temporal coherence reduces the SAXS intensity. Applying this formalism to two benchmark geometries - single density steps (wires) and periodic gratings -- we obtain analytic estimates. For gratings, finite coherence simultaneously lowers Bragg-peak heights and broadens their widths, whereas for isolated steps only the overall scaling with q affected. We map the parameter space relevant to current SASE and self-seeded XFELs, revealing that Doppler effects remain managable for the trieval of geometry parameters (less than few 10 % error) for SASE bandwidths but become the dominant error source in seeded configurations or above-keV temperatures. Practical consequences for density-gradient retrieval and interface-sharpness measurements are quantified. The results supply clear criteria for when Doppler broadening must be included in SAXS data analysis and offer a route to infer electron temperature directly from coherence-loss signatures.

physics.plasm-ph

Radiative Signatures of Magnetic Reconnection: An Approach to Remote Probing of Reconnection Dynamics

Magnetic reconnection drives a wide range of astrophysical phenomena, including geomagnetic storms, solar flares, and activity in blazars. However, direct measurement of key reconnection observables remains challenging due to the remote and extreme nature of these environments. While high-energy particle showers observed on Earth are often attributed to reconnection, the underlying mechanisms are not fully understood, and clear diagnostic signatures are lacking. We present a theoretical, data-driven approach for identifying reconnection radiation signatures and enabling remote diagnostics of reconnection in astrophysical settings through radiation spectra. Using particle-in-cell (PIC) simulations of magnetic reconnection, we generate radiation spectra and establish connections between spectral features and the underlying reconnection dynamics. We develop a method to estimate the ratio of the reconnection electric field to the plasmoid magnetic field from spectral data. Analytic calculations show that other parameters can be extracted in the ultra-relativistic reconnection regime, such as the magnetic field or the current sheet width.

astro-ph.HE

The Artificial Scientist -- in-transit Machine Learning of Plasma Simulations

Increasing HPC cluster sizes and large-scale simulations that produce petabytes of data per run, create massive IO and storage challenges for analysis. Deep learning-based techniques, in particular, make use of these amounts of domain data to extract patterns that help build scientific understanding. Here, we demonstrate a streaming workflow in which simulation data is streamed directly to a machine-learning (ML) framework, circumventing the file system bottleneck. Data is transformed in transit, asynchronously to the simulation and the training of the model. With the presented workflow, data operations can be performed in common and easy-to-use programming languages, freeing the application user from adapting the application output routines. As a proof-of-concept we consider a GPU accelerated particle-in-cell (PIConGPU) simulation of the Kelvin- Helmholtz instability (KHI). We employ experience replay to avoid catastrophic forgetting in learning from this non-steady process in a continual manner. We detail challenges addressed while porting and scaling to Frontier exascale system.

physics.comp-ph

A Parallel and Highly-Portable HPC Poisson Solver: Preconditioned Bi-CGSTAB with alpaka

This paper presents the design, implementation, and performance analysis of a parallel and GPU-accelerated Poisson solver based on the Preconditioned Bi-Conjugate Gradient Stabilized (Bi-CGSTAB) method. The implementation utilizes the MPI standard for distributed-memory parallelism, while on-node computation is handled using the alpaka framework: this ensures both shared-memory parallelism and inherent performance portability across different hardware architectures. We evaluate the solver's performances on CPUs and GPUs (NVIDIA Hopper H100 and AMD MI250X), comparing different preconditioning strategies, including Block Jacobi and Chebyshev iteration, and analyzing the performances both at single and multi-node level. The execution efficiency is characterized with a strong scaling test and using the AMD Omnitrace profiling tool. Our results indicate that a communication-free preconditioner based on the Chebyshev iteration can speed up the solver by more than six times. The solver shows comparable performances across different GPU architectures, achieving a speed-up in computation up to 50 times compared to the CPU implementation. In addition, it shows a strong scaling efficiency greater than 90% up to 64 devices.

cs.DC

Streaming Data in HPC Workflows Using ADIOS

The "IO Wall" problem, in which the gap between computation rate and data access rate grows continuously, poses significant problems to scientific workflows which have traditionally relied upon using the filesystem for intermediate storage between workflow stages. One way to avoid this problem in scientific workflows is to stream data directly from producers to consumers and avoiding storage entirely. However, the manner in which this is accomplished is key to both performance and usability. This paper presents the Sustainable Staging Transport, an approach which allows direct streaming between traditional file writers and readers with few application changes. SST is an ADIOS "engine", accessible via standard ADIOS APIs, and because ADIOS allows engines to be chosen at run-time, many existing file-oriented ADIOS workflows can utilize SST for direct application-to-application communication without any source code changes. This paper describes the design of SST and presents performance results from various applications that use SST, for feeding model training with simulation data with substantially higher bandwidth than the theoretical limits of Frontier's file system, for strong coupling of separately developed applications for multiphysics multiscale simulation, or for in situ analysis and visualization of data to complete all data processing shortly after the simulation finishes.

cs.PF

Enabling High-Throughput Parallel I/O in Particle-in-Cell Monte Carlo Simulations with openPMD and Darshan I/O Monitoring

Large-scale HPC simulations of plasma dynamics in fusion devices require efficient parallel I/O to avoid slowing down the simulation and to enable the post-processing of critical information. Such complex simulations lacking parallel I/O capabilities may encounter performance bottlenecks, hindering their effectiveness in data-intensive computing tasks. In this work, we focus on introducing and enhancing the efficiency of parallel I/O operations in Particle-in-Cell Monte Carlo simulations. We first evaluate the scalability of BIT1, a massively-parallel electrostatic PIC MC code, determining its initial write throughput capabilities and performance bottlenecks using an HPC I/O performance monitoring tool, Darshan. We design and develop an adaptor to the openPMD I/O interface that allows us to stream PIC particle and field information to I/O using the BP4 backend, aggressively optimized for I/O efficiency, including the highly efficient ADIOS2 interface. Next, we explore advanced optimization techniques such as data compression, aggregation, and Lustre file striping, achieving write throughput improvements while enhancing data storage efficiency. Finally, we analyze the enhanced high-throughput parallel I/O and storage capabilities achieved through the integration of openPMD with rapid metadata extraction in BP4 format. Our study demonstrates that the integration of openPMD and advanced I/O optimizations significantly enhances BIT1's I/O performance and storage capabilities, successfully introducing high throughput parallel I/O and surpassing the capabilities of traditional file I/O.

cs.DC

EZ: An Efficient, Charge Conserving Current Deposition Algorithm for Electromagnetic Particle-In-Cell Simulations

We present EZ, a novel current deposition algorithm for particle-in-cell (PIC) simulations. EZ calculates the current density on the electromagnetic grid due to macro-particle motion within a time step by solving the continuity equation of electrodynamics. Being a charge conserving hybridization of Esirkepov's method and ZigZag, we refer to it as ``EZ'' as shorthand for ``Esirkepov meets ZigZag''. Simulations of a warm, relativistic plasma with PIConGPU show that EZ achieves the same level of charge conservation as the commonly used method by Esirkepov, yet reaches higher performance for macro-particle assignment-functions up to third-order. In addition to a detailed description of the functioning of EZ, reasons for the expected and observed performance increase are given, and guidelines for its implementation aiming at highest performance on GPUs are provided.

physics.comp-ph

Challenges Porting a C++ Template-Metaprogramming Abstraction Layer to Directive-based Offloading

HPC systems employ a growing variety of compute accelerators with different architectures and from different vendors. Large scientific applications are required to run efficiently across these systems but need to retain a single code-base in order to not stifle development. Directive-based offloading programming models set out to provide the required portability, but, to existing codes, they themselves represent yet another API to port to. Here, we present our approach of porting the GPU-accelerated particle-in-cell code PIConGPU to OpenACC and OpenMP target by adding two new backends to its existing C++-template metaprogramming-based offloading abstraction layer alpaka and avoiding other modifications to the application code. We introduce our approach in the face of conflicts between requirements and available features in the standards as well as practical hurdles posed by immature compiler support.

cs.PL

Realizing Quantum free-electron lasers: A critical analysis of experimental challenges and theoretical limits

We examine the experimental requirements for realizing a high-gain Quantum free-electron laser (Quantum FEL). Beyond fundamental constraints on electron beam and undulator, we discuss optimized interaction geometries, include coherence properties along with the impact of diffraction, space-charge and spontaneous emission. Based on desired Quantum FEL properties, as well as current experimental capabilities, we provide a procedure for determining a corresponding set of experimental parameters. Even for an idealized situation, the combined constraints on space-charge and spontaneous emission put strong limits on sustaining the quantum regime over several gain lengths. Guided by these results we propose to shift the focus towards seeded Quantum FELs instead of continuing to aim for self-amplified spontaneous emission (SASE). Moreover, we point out the necessity of a rigorous quantum theory for spontaneous emission as well as for space-charge in order to identify possible loopholes in our line of argument.

physics.acc-ph

Quantitatively consistent computation of coherent and incoherent radiation in particle-in-cell codes - a general form factor formalism for macro-particles

Quantitative predictions from synthetic radiation diagnostics often have to consider all accelerated particles. For particle-in-cell (PIC) codes, this not only means including all macro-particles but also taking into account the discrete electron distribution associated with them. This paper presents a general form factor formalism that allows to determine the radiation from this discrete electron distribution in order to compute the coherent and incoherent radiation self-consistently. Furthermore, we discuss a memory-efficient implementation that allows PIC simulations with billions of macro-particles. The impact on the radiation spectra is demonstrated on a large scale LWFA simulation.

physics.comp-ph