SearcharxivSearch

arXiv subjects

Daniele Gregori

Publications and source records attributed to Daniele Gregori.

18 recordsLinked to original sources

Portable Acceleration of Learning With Errors KEMs for Post-Quantum Cryptography

The transition to post-quantum cryptography (PQC) is driving demand for implementations that can meet the computational requirements of real-world applications. Among the proposed PQC constructions, Learning With Errors (LWE) based key encapsulation mechanisms (KEMs) are particularly attractive due to their strong security foundations, but they incur substantial computational costs from matrix operations and large-scale cryptographically secure random number generation. These characteristics position GPU acceleration as an effective approach for lowering the computational overhead of lattice based cryptographic schemes. In this work, we present a portable GPU implementation of a plain LWE based KEM using OpenMP Target offloading. Unlike most existing GPU implementations, which rely on CUDA specific optimizations, our approach uses a single source code base that executes on both NVIDIA and AMD accelerators. We evaluate the proposed implementation on different accelerator architectures, analyzing performance benchmarking, runtime profiling, scalability analysis, and energy to solution measurements. Experimental results show that OpenMP Target offloading delivers substantial acceleration over a multicore CPU baseline while preserving source level portability across heterogeneous GPU ecosystems. Cross platform analysis identifies NVIDIA GH200 and AMD MI300X as the most effective platforms for this memory bound workload, while profiling indicates that memory system organization and CPU GPU interaction play a more critical role than peak compute capability alone. These findings demonstrate that portable GPU acceleration can significantly reduce the computational overhead of PQC while avoiding vendor lock in, thereby facilitating the deployment of quantum resistant cryptographic infrastructures.

cs.CR

Three ways to share a QPU: Scheduling strategies for hybrid Quantum-HPC applications

As quantum computing (QC) technologies mature, their integration into established high-performance computing (HPC) infrastructures is becoming a central objective for next-generation computing systems. However, unlocking the potential of hybrid platforms for computationally demanding workloads remains challenging. The mismatch between quantum and classical programming models, the limited maturity of quantum software stacks, and the scarcity of quantum processing units (QPUs) above all, necessitate scheduling strategies that go beyond standard HPC mechanisms to manage such heterogeneous and constrained resources. To address this issue, we investigate three distinct methodologies for HPC-QC resource scheduling: time-based multiplexing, dynamic resource management, and workflow decomposition. Experimental validation on production HPC clusters and real quantum hardware demonstrates the effectiveness of these approaches under different workload scenarios. Malleability and workflow strategies significantly optimize classical resource utilization, reducing consumption by up to 45.7% and 64% respectively, proving to be best fitted for hybrid jobs where quantum and classical workloads are evenly balanced. Conversely, time-multiplexing enhances QPU utilization and reduces execution time at the cluster level, making it the optimal strategy for the opposite context, which is characterized by high classical-quantum workload imbalances. These findings underscore the practical viability of tailored scheduling strategies for hybrid HPC-QC environments and highlight their complementarity in building efficient, scalable software stacks for next-generation quantum-accelerated facilities.

quant-ph

GPU Acceleration of Learning With Errors KEMs Using OpenACC for Post-Quantum Cryptography

Shor's algorithm proved that asymmetric cryptographic protocols based on the integer factorization and discrete logarithm problems are no longer safe in a world with large-scale quantum computers. As a result, Post-Quantum Cryptography (PQC) has been developed over the last few years, seeking cryptographic primitives resistant to quantum attacks. One of the main hard problems underlying PQC schemes is the Learning with Errors (LWE) problem, which is significantly more computationally intensive than its classical predecessors. In this work, we present a Key Encapsulation Mechanism (KEM) based on plain LWE and develop a GPU-oriented implementation using OpenACC. We evaluate the performance of our accelerated application in terms of both time-to-solution and energy-to-solution, considering bare-metal and containerized executions across multiple NVIDIA GPU models and generations. Our implementation achieves significant acceleration across all tested GPU platforms. In particular, on the NVIDIA Grace Hopper Superchip, it attains up to a $208\times$ speedup over a multithreaded CPU baseline and enables the execution of problem sizes that are impractical on CPU architectures due to memory and synchronization constraints. Energy consumption analysis also shows $\approx 2\times$ better efficiency when using the Superchip compared to systems equipped with x86-based CPUs and NVIDIA H100 GPUs. These results highlight the effectiveness of GPU acceleration for computationally demanding LWE-based cryptographic workloads.

cs.CR

Assessing Performance and Porting Strategies for Gravitational $N$-Body Simulations on the RISC-V-Based Tenstorrent Wormhole\textsuperscript{\texttrademark}

While RISC-V-based accelerators were initially designed with artificial intelligence applications in mind, they are increasingly being recognized as promising platforms for high performance scientific computing. In this work, we present three strategies for scaling an $N$-body code across multiple Tenstorrent Wormhole accelerators based on the RISC-V architecture. We assess the performance of these approaches by measuring both the execution time and the energy consumption required to complete a representative simulation, ultimately identifying the configuration that offers the most favorable balance between efficiency and performance.

cs.DC

A Physically-Informed Subgraph Isomorphism Approach to Molecular Docking Using Quantum Annealers

Molecular docking is a crucial step in the development of new drugs as it guides the positioning of a small molecule (ligand) within the pocket of a target protein. In the literature, a feasibility study explored the potential of D-Wave quantum annealers for purely geometric molecular docking, neglecting physicochemical interactions between the protein and the ligand and focusing solely on their simplified geometries. To achieve this, the ligands were represented as graphs incorporating their geometric properties and then mapped onto a grid that discretized the three-dimensional space of the protein pocket. The quality of the ligand pose on the protein pocket was evaluated through the isomorphism between the ligand graph and the spatial grid. This paper builds on the previous study by introducing physicochemical interactions between the protein-ligand pair into the QUBO problem to improve the accuracy of the docking results. This paper presents a novel QUBO formulation that includes Coulomb and van der Waals forces, together with components representing H-bond and hydrophobic interactions. We integrate these physical interactions as corrective terms to the previous purely geometric QUBO formulation, and provide experimental results using the D-Wave quantum annealers to demonstrate their impact on the accuracy of the docking results.

cs.ET

EuroHPC SPACE CoE: Redesigning Scalable Parallel Astrophysical Codes for Exascale

High Performance Computing (HPC) based simulations are crucial in Astrophysics and Cosmology (A&C), helping scientists investigate and understand complex astrophysical phenomena. Taking advantage of exascale computing capabilities is essential for these efforts. However, the unprecedented architectural complexity of exascale systems impacts legacy codes. The SPACE Centre of Excellence (CoE) aims to re-engineer key astrophysical codes to tackle new computational challenges by adopting innovative programming paradigms and software (SW) solutions. SPACE brings together scientists, code developers, HPC experts, hardware (HW) manufacturers, and SW developers. This collaboration enhances exascale A&C applications, promoting the use of exascale and post-exascale computing capabilities. Additionally, SPACE addresses high-performance data analysis for the massive data outputs from exascale simulations and modern observations, using machine learning (ML) and visualisation tools. The project facilitates application deployment across platforms by focusing on code repositories and data sharing, integrating European astrophysical communities around exascale computing with standardised SW and data protocols.

astro-ph.IM

Towards Exascale Computing for Astrophysical Simulation Leveraging the Leonardo EuroHPC System

Developing and redesigning astrophysical, cosmological, and space plasma numerical codes for existing and next-generation accelerators is critical for enabling large-scale simulations. To address these challenges, the SPACE Center of Excellence (SPACE-CoE) fosters collaboration between scientists, code developers, and high-performance computing experts to optimize applications for the exascale era. This paper presents our strategy and initial results on the Leonardo system at CINECA for three flagship codes, namely gPLUTO, OpenGadget3 and iPIC3D, using profiling tools to analyze performance on single and multiple nodes. Preliminary tests show all three codes scale efficiently, reaching 80% scalability up to 1,024 GPUs.

cs.DC

Integrability, susy $SU(2)$ matter gauge theories and black holes

We show that previous correspondence between some (integrable) statistical field theory quantities and periods of $SU(2)$ $\mathcal{N}=2$ deformed gauge theory still holds if we add $N_f=1,2$ flavours of matter. Moreover, the correspondence entails a new non-perturbative solution to the theory. Eventually, we use this solution to give exact results on quasinormal modes of black branes and holes.

hep-th

Accelerating Gravitational $N$-Body Simulations Using the RISC-V-Based Tenstorrent Wormhole

Although originally developed primarily for artificial intelligence workloads, RISC-V-based accelerators are also emerging as attractive platforms for high-performance scientific computing. In this work, we present our approach to accelerating an astrophysical $N$-body code on the RISC-V-based Wormhole n300 card developed by Tenstorrent. Our results show that this platform can be highly competitive for astrophysical simulations employing this class of algorithms, delivering more than a $2 \times$ speedup and approximately $2 \times$ energy savings compared to a highly optimized CPU implementation of the same code.

cs.DC

Dynamic Solutions for Hybrid Quantum-HPC Resource Allocation

The integration of quantum computers within classical High-Performance Computing (HPC) infrastructures is receiving increasing attention, with the former expected to serve as accelerators for specific computational tasks. However, combining HPC and quantum computers presents significant technical challenges, including resource allocation. This paper presents a novel malleability-based approach, alongside a workflow-based strategy, to optimize resource utilization in hybrid HPC-quantum workloads. With both these approaches, we can release classical resources when computations are offloaded to the quantum computer and reallocate them once quantum processing is complete. Our experiments with a hybrid HPC-quantum use case show the benefits of dynamic allocation, highlighting the potential of those solutions.

quant-ph

Monte Cimone v2: Down the Road of RISC-V High-Performance Computers

Many RISC-V (RV) platforms and SoCs have been announced in recent years targeting the HPC sector, but only a few of them are commercially available and engineered to fit the HPC requirements. The Monte Cimone project targeted assessing their capabilities and maturity, aiming to make RISC-V a competitive choice when building a datacenter. Nowadays, Systems-on-chip (SoCs) featuring RV cores with vector extension, form factor and memory capacity suitable for HPC applications are available in the market, but it is unclear how compilers and open-source libraries can take advantage of its performance. In this paper, we describe the performance assessment of the upgrade of the Monte Cimone (MCv2) cluster with the Sophgo SG2042 processor on HPC workloads. Also adding an exploration of BLAS libraries optimization. The upgrade increases the attained node's performance by 127x on HPL DP FLOP/s and 69x on Stream Memory Bandwidth.

cs.DC

Assessing the Elephant in the Room in Scheduling for Current Hybrid HPC-QC Clusters

Quantum computing resources are among the most promising candidates for extending the computational capabilities of High-Performance Computing (HPC) systems. As a result, HPC-quantum integration has become an increasingly active area of research. While much of the existing literature has focused on software stack integration and quantum circuit compilation, key challenges such as hybrid resource allocation and job scheduling-especially relevant in the current Noisy Intermediate-Scale Quantum era-have received less attention. In this work, we highlight these critical issues in the context of integrating quantum computers with operational HPC environments, taking into account the current maturity and heterogeneity of quantum technologies. We then propose a set of conceptual strategies aimed at addressing these challenges and paving the way for practical HPC-QC integration in the near future.

quant-ph

To Repair or Not to Repair: Assessing Fault Resilience in MPI Stencil Applications

With the increasing size of HPC computations, faults are becoming more and more relevant in the HPC field. The MPI standard does not define the application behaviour after a fault, leaving the burden of fault management to the user, who usually resorts to checkpoint and restart mechanisms. This trend is especially true in stencil applications, as their regular pattern simplifies the selection of checkpoint locations. However, checkpoint and restart mechanisms introduce non-negligible overhead, disk load, and scalability concerns. In this paper, we show an alternative through fault resilience, enabled by the features provided by the User Level Fault Mitigation extension and shipped within the Legio fault resilience framework. Through fault resilience, we continue executing only the non-failed processes, thus sacrificing result accuracy for faster fault recovery. Our experiments on a specimen stencil application show that, despite the fault impact visible in the result, we produced meaningful values usable for scientific research, proving the possibilities of a fault resilience approach in a stencil scenario.

cs.DC

Tunable and Portable Extreme-Scale Drug Discovery Platform at Exascale: the LIGATE Approach

Today digital revolution is having a dramatic impact on the pharmaceutical industry and the entire healthcare system. The implementation of machine learning, extreme-scale computer simulations, and big data analytics in the drug design and development process offers an excellent opportunity to lower the risk of investment and reduce the time to the patient. Within the LIGATE project, we aim to integrate, extend, and co-design best-in-class European components to design Computer-Aided Drug Design (CADD) solutions exploiting today's high-end supercomputers and tomorrow's Exascale resources, fostering European competitiveness in the field. The proposed LIGATE solution is a fully integrated workflow that enables to deliver the result of a virtual screening campaign for drug discovery with the highest speed along with the highest accuracy. The full automation of the solution and the possibility to run it on multiple supercomputing centers at once permit to run an extreme scale in silico drug discovery campaign in few days to respond promptly for example to a worldwide pandemic crisis.

cs.DC

Fault Awareness in the MPI 4.0 Session Model

The latest version of MPI introduces new functionalities like the Session model, but it still lacks fault management mechanisms. Past efforts produced tools and MPI standard extensions to manage fault presence, including ULFM. These measures are effective against faults but do not fully support the new additions to the standard. In this paper, we combine the fault management possibilities of ULFM with the new Session model functionality introduced in version 4.0 of the standard. We focus on the communicator creation procedure, highlighting criticalities and proposing a method to circumvent them. The experimental campaign shows that the proposed solution does not significantly affect applications' execution time and scalability while better managing the insurgence of faults.

cs.DC

Monte Cimone: Paving the Road for the First Generation of RISC-V High-Performance Computers

The new open and royalty-free RISC-V ISA is attracting interest across the whole computing continuum, from microcontrollers to supercomputers. High-performance RISC-V processors and accelerators have been announced, but RISC-V-based HPC systems will need a holistic co-design effort, spanning memory, storage hierarchy interconnects and full software stack. In this paper, we describe Monte Cimone, a fully-operational multi-blade computer prototype and hardware-software test-bed based on U740, a double-precision capable multi-core, 64-bit RISC-V SoC. Monte Cimone does not aim to achieve strong floating-point performance, but it was built with the purpose of "priming the pipe" and exploring the challenges of integrating a multi-node RISC-V cluster capable of providing an HPC production stack including interconnect, storage and power monitoring infrastructure on RISC-V hardware. We present the results of our hardware/software integration effort, which demonstrate a remarkable level of software and hardware readiness and maturity - showing that the first generation of RISC-V HPC machines may not be so far in the future.

cs.DC

A new method for exact results on Quasinormal Modes of Black Holes

We develop a new method for writing simple exact equations characterizing gravity solutions among which black holes and in particular the quasinormal modes. More precisely, we derive the full system of functional and Thermodynamic Bethe Ansatz non linear integral equations of quantum integrability. In particular, we prove that the Quasinormal Modes verify different equivalent exact quantization conditions and identify them with Bethe roots. We numerically solve the integral equation and compare the results with other methods. Eventually, we can definitely certify its simplicity, accuracy and effectiveness. Furthermore, this method connects different unexpected fields and paves the way for innovative ways of investigations in gravity and gauge theories.

hep-th

Integrability and cycles of deformed ${\cal N}=2$ gauge theory

To analyse pure ${\cal N}=2$ $SU(2)$ gauge theory in the Nekrasov-Shatashvili (NS) limit (or deformed Seiberg-Witten (SW)), we use the Ordinary Differential Equation/Integrable Model (ODE/IM) correspondence, and in particular its (broken) discrete symmetry in its extended version with {\it two} singular irregular points. Actually, this symmetry appears to be 'manifestation' of the spontaneously broken $\mathbb{Z}_2$ R-symmetry of the original gauge problem and the two deformed SW cycles are simply connected to the Baxter's $T$ and $Q$ functions, respectively, of the Liouville conformal field theory at the self-dual point. The liaison is realised via a second order differential operator which is essentially the 'quantum' version of the square of the SW differential. Moreover, the constraints imposed by the broken $\mathbb{Z}_2$ R-symmetry acting on the moduli space (Bilal-Ferrari equations) seem to have their quantum counterpart in the $TQ$ and the $T$ periodicity relations, and integrability yields also a useful Thermodynamic Bethe Ansatz (TBA) for the cycles ($Y(θ,\pm u)$ or their square roots, $Q(θ,\pm u)$). A latere, two efficient asymptotic expansion techniques are presented. Clearly, the whole construction is extendable to gauge theories with matter and/or higher rank groups.

hep-th