SearcharxivSearch

arXiv subjects

Stefano Markidis

Publications and source records attributed to Stefano Markidis.

At least 19 recordsLinked to original sources

Computational Methods and GPU Acceleration in Plasma Physics: A Empirical Analysis of arXiv Publications and Research Trends

Computational plasma physics increasingly relies on high-performance computing (HPC) methods, including particle-in-cell (PIC), gyrokinetic, and magnetohydrodynamic (MHD) simulations, yet field-wide evidence on how method choice, team size, and GPU adoption shape research outputs remains limited. We analyze 5,522 computational plasma physics papers published on arXiv between 2010 and 2025 using large-scale text mining, employing abstract length as a proxy for methodological and algorithmic complexity. Using ordinary least squares (OLS), tobit, and logistic regression models, we examine how abstract length and GPU mentions vary with computational method, number of authors, and publication year. Controlling for collaboration size and temporal trends, results show that MHD studies have longer abstracts than PIC and gyrokinetic papers, indicating more extensive methodological and physical exposition. Abstract length increases modestly with team size, while temporal effects suggest gradual changes in abstract conciseness over time. At the same time, PIC and gyrokinetic methods have grown substantially in relative prevalence over the past decade and are strongly linked to GPU adoption, reflecting their higher computational intensity and suitability for accelerator-based architectures. Together, these findings highlight a decoupling between methodological verbosity and method prevalence, offering a new bibliometric perspective on the evolution of HPC-driven plasma physics research.

physics.plasm-ph

Fewer Histories, Faster Paths: Distributed Quantum Circuit Feynman Simulation via History Reduction, Checkpointing, and Pruning

We present a distributed method for exact sparse-output quantum circuit simulation based on the pure Feynman sum-over-histories formulation. The method computes selected computational-basis amplitudes exactly and addresses the exponential growth of the path sum through a reduced history formulation based on internal-wire assignments, determinism propagation, artificial sources, pruning, and checkpointed reuse. Boundary constraints are propagated through deterministic and wire-preserving gates, and explicit branching variables are introduced only where residual ambiguity remains. Shared work across related histories is captured via an autotuned checkpointed partition. The parallel execution model combines decomposition over requested outputs with concurrent history evaluation, while a dynamic server-worker architecture mitigates load imbalance from irregular branching and pruning. Across the circuit families studied, the method adapts to different structural regimes of the reduced history space: zero artificial sources for QFT under backward analysis, substantial speedups from checkpointing and autotuning for amplitude amplification, and a runtime-fidelity tradeoff from threshold pruning for QAOA. On quantum walk circuits, it reconstructs exact selected-output distributions up to 100 qubits and achieves 85% parallel efficiency on 8,192 CPU cores of a supercomputer.

cs.ET

High-Performance Resilient Multi-GPU Hybrid Particle-in-Cell Monte Carlo Simulations at Scale

The increasing demand for high-performance computing in plasma physics has driven scalable and resilient simulation methods capable of efficiently exploiting modern multi-GPU architectures. This work extends a portable hybrid MPI+OpenMP implementation of BIT1, focusing on high-performance resilience for accelerated Particle-in-Cell (PIC) Monte Carlo (MC) simulations under both uniform and non-uniform load conditions. Scalable particle load balancing and robust checkpoint/restart mechanisms across Nvidia and AMD accelerators are integrated with standardized I/O using openPMD and ADIOS2. This leverages BP4 for high-performance file-based checkpointing and SST for in-memory data streaming, enabling efficient data movement, resilient large-scale execution, seamless continuation from existing checkpoints, and effective handling of computational and I/O workloads. Advanced HPC profiling and tracing tools, including Nvidia Nsight Systems and AMD ROC-Profiler with Perfetto, provide detailed insights into computation, communication, and system-level behavior for optimization. Performance results on Frontier (OLCF-5), MN5, and LUMI-G demonstrate strong and weak scaling up to 800 GPUs, validating the framework for large-scale PIC MC simulations, while in-situ analysis and visualization using scalable I/O further enhance scientific insight without interrupting multi-GPU execution on current and future exascale systems.

physics.plasm-ph

Quantum Kernels are Spectral Tensor Networks

Quantum kernels admit Fourier representations whose frequencies are determined by the data-encoding gates of the underlying feature map. We show that entangling tensor kernels are matrix product operator factorizations of the corresponding Fourier coefficient tensors, thereby identifying quantum kernels as spectral tensor networks. By grouping gate-level frequency configurations that yield the same feature-wise frequency, we obtain a grouped Fourier form that induces a more compact spectral tensor network representation of the kernel. We further show that kernel target alignment serves as a bridge between the Fourier and tensor network views. On a grid that resolves the accessible Fourier modes, it becomes the Frobenius cosine similarity between Fourier coefficient tensors. Our numerical experiments show that layered quantum kernels admit accurate representations with small bond dimension, revealing a compressibility governed by correlations between Fourier modes. This compressibility provides a diagnostic of classical representability and of whether kernel evaluation is likely to remain classically tractable.

quant-ph

Not Your Usual FFT: QFT$\rightarrow$FFT via Classical Quantum-Circuit Simulation

We introduce QFT$\rightarrow$FFT, a family of HPC FFT libraries that compute the discrete Fourier transform by executing a quantum Fourier transform (QFT) circuit on classical quantum computer simulators. Input arrays are mapped directly to state amplitudes with explicit normalization/indexing, making QFT a drop-in replacement for FFT primitives. A backend-agnostic planner builds a fused-gate schedule and memory layout adapters to increase arithmetic intensity and reduce memory data movement. We implement this design on top of Google's C++ \texttt{qsim} and evaluate OpenMP, AVX, and CUDA backends. On an AMD EPYC Zen2 processor, our AVX performance is on par with that of multithreaded FFTW, utilizing 64 threads. On an NVIDIA A100, the CUDA backend achieves more than $4\times$ lower time than both AVX and FFTW on AMD EPYC Zen2 at larger sizes. We also employ an approximate QFT (AQFT) that truncates small-angle controlled rotations beyond a cutoff $k$, reducing circuit depth and runtime while preserving accuracy.

cs.ET

Which Superconducting Qubit Model is Good Enough? From Effective Two-Level to Circuit-Based Hamiltonians for Pulse-Level Simulation

Pulse-level simulators are the lowest-level, most widely used abstraction layer for studying how quantum hardware responds to control signals, but they can be built on Hamiltonian models with very different fidelity and cost. This raises the question: which level of physical abstraction is sufficient for a given simulation objective? We study this question for a flux-tunable two-qubit superconducting device with a fixed bus coupler by comparing three Hamiltonian descriptions of the same hardware: an effective two-level model, a three-mode Duffing model, and a circuit-based transmon model in the charge basis. Using a realistic parameter set, we evaluate these models on a common benchmark suite spanning flux-dependent spectra, extracted two-qubit interaction terms, driven single-qubit dynamics, CZ gate dynamics, leakage outside the computational subspace, and runtime. Across the tested flux range, the Duffing model follows the circuit-based reference more closely than the effective model for static spectra and reduced two-qubit quantities, while in driven benchmarks, the multilevel models reveal effects absent in the effective description. Overall, the results support a layered use of abstraction in pulse-level simulation: effective models for reduced analyses, Duffing models as a practical multilevel default, and circuit-based models for high-fidelity reference simulation or detailed leakage analysis.

cs.ET

When Noisy Quantum Order Finding Remains Recoverable for Shor's Algorithm

Order finding is the core subroutine of Shor's algorithm. On NISQ hardware, phase estimation output distributions are often distorted by noise, making correct order recovery difficult. We study recoverability in noisy order finding: given a measured precision-register distribution, when does standard classical post-processing still return the true order? We analyze 680 distributions from IBM quantum systems across problem instances and circuit settings. For each distribution, we apply continued-fraction post-processing with modular verification and define recoverability as whether the recovered order equals the true one. We characterize each distribution using four features: autocorrelation peak strength, normalized entropy, dominant verified mass fraction, and verified margin fraction. We evaluate these quantities using marginal feature comparisons, single-feature AUROC analysis, and multivariate tree-based classifiers. We use random-forest permutation importance to assess which quantities contribute distinct predictive information once the other features are known. To make classification behavior interpretable, we train a decision tree that exposes threshold rules for recoverable and non-recoverable distributions. We find that recoverability is strongly associated with residual comb-like structure in the measured distribution and the way verified probability mass is organized across candidate denominators. The dominant verified mass fraction is the strongest single-feature indicator of recoverability, and tree-based analysis shows it also provides the primary split in an interpretable threshold description. Some highly distorted distributions remain recoverable when one verified denominator dominates the post-processing mass, while some visibly structured distributions fail because classical post-processing favors an incorrect verified denominator.

quant-ph

Robust Matrix-Free Newton-Krylov Solvers via Automatic Differentiation

Jacobian-Free Newton-Krylov (JFNK) methods avoid forming the full Jacobian, but still require Jacobian-vector products, i.e., Gateaux derivatives of the nonlinear residual along Krylov directions. In standard Finite Differences (FD) formulations, these products are obtained by perturbing the Newton state and differencing residuals, making the linearization sensitive to round-off error and floating-point precision. This work evaluates the global impact of forward-mode Automatic Differentiation (AD) as a replacement for FD Jacobian-vector product in finite-precision JFNK solvers. The comparison keeps the discretization, Newton iteration, line search, Krylov methods, tolerances, and CPU/GPU backend fixed, only varying linearization strategy. Benchmarks include Burgers dynamics, Su-Olson radiation diffusion, reaction-diffusion, and nonlinear time-harmonic Maxwell equations, each evaluated in different nonlinear regimes. By preventing degradation of the Krylov operator, AD accelerates computation by 2-3 orders of magnitude across both CPU and GPU architectures. More importantly, it drastically improves global solver robustness, achieving a minimum completion rate of 95%, compared to just 42% for FD. Ultimately, accurate Gateaux derivatives unify performance and accuracy in JFNK methods, making AD the optimal choice for stiff nonlinear and reduced-precision environments.

cs.CE

Post-Moore Technologies for Plasma Simulation: A Community Roadmap

Plasma simulations are among the most computationally demanding scientific workloads, combining high-dimensional kinetic evolution, particle-mesh coupling, field solves, and data-intensive communication. As general-purpose processor scaling slows, post-Moore technologies are being explored to address bottlenecks in data movement, memory access, and power consumption. This paper provides a community perspective on the role of these technologies in plasma simulation, assessing three major classes: reconfigurable and data-path accelerators, non-von Neumann architectures, and quantum computing. Each is evaluated, in a co-design approach, against representative plasma workloads spanning particle-in-cell, continuum Vlasov, gyrokinetic, fluid/MHD, hybrid, and warm dense matter methods. We find that no single technology can replace existing HPC platforms. Instead, three tiers of opportunity emerge: FPGA-class and data-path accelerators offer near-term kernel offload and workflow-level data services, non-von Neumann architectures represent medium-term directions for operator-level acceleration, and quantum computing, although the least mature, is potentially the most disruptive for warm dense matter and inertial confinement fusion microphysics. We outline best practices for selective adoption and identify focused demonstrators, benchmarking, and modular software ecosystems as immediate community priorities.

cs.ET

Quantum Optimization for Electromagnetics: Physics-Informed QAOA for Reconfigurable Intelligent Surfaces

Optimizing Reconfigurable Intelligent Surfaces (RIS) is a high-dimensional combinatorial challenge. Current quantum algorithms often simplify this problem by ignoring physical constraints like mutual coupling, which significantly degrades real-world performance. Rather than targeting a fully realistic RIS description, we embed progressively more physics-informed models of mutual coupling into Quadratic Unconstrained Binary Optimization (QUBO) formulations. We evaluate four Ising interaction models ($J_{ij}$) for the Quantum Approximate Optimization Algorithm (QAOA), ranging from idealized phase-only to fully dense physical models. Analyzing a $5 \times 5$ grid, our results expose a critical trade-off between spatial pointing accuracy and quantum hardware feasibility. While complete global coupling maximizes beamforming precision, dense Hamiltonians introduce prohibitive routing overhead and complicate convergence on near-term processors. Ultimately, we demonstrate that while physics-informed quantum optimization is mathematically viable, sparse, distance-penalized models remain a necessary compromise for execution on current noisy intermediate-scale quantum (NISQ) devices.

cs.CE

Mass Matrix Assembly on Tensor Cores for Implicit Particle-In-Cell Methods

Matrix-multiply-accumulate (MMA) units, or tensor cores, are now widespread across modern computing architectures. Yet, their use for particle-grid operators remains limited. In implicit particle methods, mass-matrix assembly is a reduction-dominated kernel in which weighted outer products of interpolation weights are accumulated over particle support. We show that this operation can be reformulated exactly, cell by cell, as a sequence of matrix products matched to hardware MMA tiles. The formulation is general with respect to interpolation order and hardware platform, and applies to both scalar mass matrices and the tensorial block mass matrix arising in implicit in the Energy-Conserving Semi-Implicit Method (ECSIM) for Particle-in-Cell simulations. We introduce particle batching and a support-group decomposition for higher-order shape functions whose stencil extends beyond a single cell, specialize the method to first- and second-order B-spline interpolation, and implement it on NVIDIA tensor cores. The resulting kernels achieve up to 3x over optimized conventional implementations and reduce end-to-end ECSIM runtime by 15%.

cs.CE

BVH-Accelerated Ray Tracing for High-Frequency Electromagnetic Backscattering

As computational complexity in electromagnetics increases with frequency, full-wave solvers become computationally infeasible for electrically large problems. To address this limitation, we present a shooting and bouncing rays (SBR) method for efficiently modeling electromagnetic backscattering of metallic objects in the high-frequency regime. The method couples multi-reflection geometrical-optics ray transport with a physical optics surface integral discretized over ray tubes. To reduce the massive ray-surface intersection search space, we use a bounding volume hierarchy (BVH) and organize the computation as a trace-integrate pipeline. The ray tracing generates hit data, and the physical optics integral is evaluated over valid intersections only. Numerical accuracy is controlled through an incident-ray sampling rule that mitigates phase aliasing in the discretized physical optics integration. The method is accelerated on NVIDIA and AMD GPUs and parallelized with MPI. We validate against analytical Mie solutions for a perfectly electrically conducting (PEC) sphere and demonstrate applicability to a complex aircraft geometry for monostatic radar cross-section prediction.

cs.CE

Making Room for AI: Multi-GPU Molecular Dynamics with Deep Potentials in GROMACS

GROMACS is a de-facto standard for classical Molecular Dynamics (MD). The rise of AI-driven interatomic potentials that pursue near-quantum accuracy at MD throughput now poses a significant challenge: embedding neural-network inference into multi-GPU simulations retaining high-performance. In this work, we integrate the MLIP framework DeePMD-kit into GROMACS, enabling domain-decomposed, GPU-accelerated inference across multi-node systems. We extend the GROMACS NNPot interface with a DeePMD backend, and we introduce a domain decomposition layer decoupled from the main simulation. The inference is executed concurrently on all processes, with two MPI collectives used each step to broadcast coordinates and to aggregate and redistribute forces. We train an in-house DPA-1 model (1.6 M parameters) on a dataset of solvated protein fragments. We validate the implementation on a small protein system, then we benchmark the GROMACS-DeePMD integration with a 15,668 atom protein on NVIDIA A100 and AMD MI250x GPUs up to 32 devices. Strong-scaling efficiency reaches 66% at 16 devices and 40% at 32; weak-scaling efficiency is 80% to 16 devices and reaches 48% (MI250x) and 40% (A100) at 32 devices. Profiling with the ROCm System profiler shows that >90% of the wall time is spent in DeePMD inference, while MPI collectives contribute <10%, primarily since they act as a global synchronization point. The principal bottlenecks are the irreducible ghost-atom cost set by the cutoff radius, confirmed by a simple throughput model, and load imbalance across ranks. These results demonstrate that production MD with near ab initio fidelity is feasible at scale in GROMACS.

cs.DC

Multi-GPU Hybrid Particle-in-Cell Monte Carlo Simulations for Exascale Computing Systems

Particle-in-Cell (PIC) Monte Carlo (MC) simulations are central to plasma physics but face increasing challenges on heterogeneous HPC systems due to excessive data movement, synchronization overheads, and inefficient utilization of multiple accelerators. In this work, we present a portable, multi-GPU hybrid MPI+OpenMP implementation of BIT1 that enables scalable execution on both Nvidia and AMD accelerators through OpenMP target tasks with explicit dependencies to overlap computation and communication across devices. Portability is achieved through persistent device-resident memory, an optimized contiguous one-dimensional data layout, and a transition from unified to pinned host memory to improve large data-transfer efficiency, together with GPU Direct Memory Access (DMA) and runtime interoperability for direct device-pointer access. Standardized and scalable I/O is provided using openPMD and ADIOS2, supporting high-performance file I/O, in-memory data streaming, and in-situ analysis and visualization. Performance results on pre-exascale and exascale systems, including Frontier (OLCF-5) for up to 16,000 GPUs, demonstrate significant improvements in run time, scalability, and resource utilization for large-scale PIC MC simulations.

physics.plasm-ph

A Continuous-Variable Quantum Fourier Layer: Applications to Filtering and PDE Solving

Fourier representations play a central role in operator learning methods for partial differential equations and are increasingly being explored in quantum machine learning architectures. The classical fast Fourier transform (FFT), particularly in its Cooley--Tukey decomposition, exhibits a structure that naturally matches continuous-variable quantum circuits. This correspondence establishes a direct structural isomorphism between the Cooley-Tukey butterfly network and Gaussian photonic gates, enabling the FFT to be realized as a native optical computation in continuous-variable quantum computing. Building on this observation, we introduce a continuous-variable Quantum Fourier Layer (CV--QFL) based on a bipartite Gaussian encoding and a Cooley-Tukey quantum Fourier transform, enabling exact two-dimensional spectral processing within a Gaussian photonic circuit. We test the CV--QFL on two representative tasks: spectral low-pass filtering and Fourier-domain integration of the heat equation. In both cases, the results match the classical reference to machine precision. Beyond these examples, our method naturally extends to optical-input settings in which the signal is already available as a Gaussian optical field. In such scenarios, coherent light coupled into single-mode waveguides can be processed directly by the CV--QFL, bypassing the need for an explicit classical-to-quantum encoding stage. This enables native spectral processing of light and lays the groundwork for new approaches to quantum scientific machine learning, in particular for future neural operator architectures within the CV framework.

quant-ph

Tensor Network Compression for Fully Spectral Vlasov-Poisson Simulation

We propose a numerical method for kinetic plasma simulation in which the phase-space distribution is represented by a low-rank tensor network with an adaptive level of compression. The Vlasov-Poisson system is advanced using a second-order Strang splitting scheme, with the advection and acceleration steps treated spectrally in position and velocity, respectively. By representing both the state and the operators required for time evolution as compressed tensor objects, the propagation can be carried out directly in tensor form without reconstructing the full phase-space grid. The self-consistent electric field is likewise obtained entirely within the tensor formalism through a tensor-based Poisson solver. We validate the approach on Landau damping and the two-stream instability, and find near-indistinguishable agreement with the corresponding full-grid reference while achieving substantial compression of the state representation. Additionally, we examine how varying the degree of compression influences conservation properties, positivity behavior, and computational cost, and show that compressibility varies strongly across the considered dynamical regimes. More broadly, the results point toward a promising class of spectral Eulerian plasma solvers that operate directly on compressed tensor representations.

physics.comp-ph

Enabling AI Deep Potentials for Ab Initio-quality Molecular Dynamics Simulations in GROMACS

State-of-the-art AI deep potentials provide ab initio-quality results, but at a fraction of the computational cost of first-principles quantum mechanical calculations, such as density functional theory. In this work, we bring AI deep potentials into GROMACS, a production-level Molecular Dynamics (MD) code, by integrating with DeePMD-kit that provides domain-specific deep learning (DL) models of interatomic potential energy and force fields. In particular, we enable AI deep potentials inference across multiple DP model families and DL backends by coupling GROMACS Neural Network Potentials with the C++/CUDA backend in DeePMD-kit. We evaluate two recent large-atom-model architectures, DPA2 that is based on the attention mechanism and DPA3 that is based on GNN, in GROMACS using four ab initio-quality protein-in-water benchmarks (1YRF, 1UBQ, 3LZM, 2PTC) on NVIDIA A100 and GH200 GPUs. Our results show that DPA2 delivers up to 4.23x and 3.18x higher throughput than DPA3 on A100 and GH200 GPUs, respectively. We also provide a characterization study to further contrast DPA2 and DPA3 in throughput, memory usage, and kernel-level execution on GPUs. Our findings identify kernel-launch overhead and domain-decomposed inference as the main optimization priorities for AI deep potentials in production MD simulations.

cs.DC

Unsupervised Physics-Informed Operator Learning through Multi-Stage Curriculum Training

Solving partial differential equations remains a central challenge in scientific machine learning. Neural operators offer a promising route by learning mappings between function spaces and enabling resolution-independent inference, yet they typically require supervised data. Physics-informed neural networks address this limitation through unsupervised training with physical constraints but often suffer from unstable convergence and limited generalization capability. To overcome these issues, we introduce a multi-stage physics-informed training strategy that achieves convergence by progressively enforcing boundary conditions in the loss landscape and subsequently incorporating interior residuals. At each stage the optimizer is re-initialized, acting as a continuation mechanism that restores stability and prevents gradient stagnation. We further propose the Physics-Informed Spline Fourier Neural Operator (PhIS-FNO), combining Fourier layers with Hermite spline kernels for smooth residual evaluation. Across canonical benchmarks, PhIS-FNO attains a level of accuracy comparable to that of supervised learning, using labeled information only along a narrow boundary region, establishing staged, spline-based optimization as a robust paradigm for physics-informed operator learning.

cs.LG