SearcharxivSearch

arXiv subjects

Maurizio Palesi

Publications and source records attributed to Maurizio Palesi.

At least 19 recordsLinked to original sources

A Survey on Design Methodologies for Accelerating Deep Learning on Heterogeneous Architectures

Given their increasing size and complexity, the need for efficient execution of deep neural networks has become increasingly pressing in the design of heterogeneous High-Performance Computing (HPC) and edge platforms, leading to a wide variety of proposals for specialized deep learning architectures and hardware accelerators. The design of such architectures and accelerators requires a multidisciplinary approach combining expertise from several areas, from machine learning to computer architecture, low-level hardware design, and approximate computing. Several methodologies and tools have been proposed to improve the process of designing accelerators for deep learning, aimed at maximizing parallelism and minimizing data movement to achieve high performance and energy efficiency. This paper critically reviews influential tools and design methodologies for Deep Learning accelerators, offering a wide perspective in this rapidly evolving field. This work complements surveys on architectures and accelerators by covering hardware-software co-design, automated synthesis, domain-specific compilers, design space exploration, modeling, and simulation, providing insights into technical challenges and open research directions.

cs.AR

COSMA: Communication-aware Optimization of Fermionic Simulation Kernels for Modular Quantum Architectures

Quantum simulation is a leading application of quantum computing, but scaling to chemically relevant problems requires modular architectures composed of interconnected quantum processing units. In such systems, inter-core quantum communication becomes a major performance bottleneck. In this work, we present COSMA, a communication-aware compilation framework for fermionic simulation kernels targeting modular quantum architectures. Our approach jointly optimizes fermion-to-qubit mapping, Pauli scheduling, and qubit allocation to minimize inter-core state transfers. Evaluated on molecular benchmarks, COSMA achieves up to $2.5\times$ reduction in communication cost compared to state-of-the-art baselines, with a median improvement of $1.7\times$. These results demonstrate that cross-layer co-design is essential for efficient and scalable quantum simulation on multi-core quantum hardware.

quant-ph

Dependency-Aware Circuit Scheduling for Multi-Core Quantum Systems to Minimize Makespan

Multi-core quantum computing architectures have emerged as a promising solution to the qubit scalability limitations of monolithic NISQ devices. Quantum algorithms are expressed as quantum circuits composed of single- and two-qubit gates. However, circuit scheduling in multi-core quantum systems remains largely unexplored. Reducing overall execution time (makespan), increasing core utilization, and hiding communication latency behind computation depends on effective scheduling. In this paper, we first introduce a layered scheduling approach as a baseline where quantum gates within the same layer are executed in parallel, while layers themselves are executed sequentially. We then propose a greedy scheduling strategy which schedules each gate as soon as all its dependencies and required resources are available. This allows fine-grained parallelism across cores. Our evaluation shows that on real benchmarks, greedy scheduling achieves an average 40% reduction in makespan and improvement in core utilization. The results suggest that the use of intelligent circuit scheduling to exploit parallelism can greatly enhance the speed of circuit execution in multi-core quantum architectures.

quant-ph

MATCHA: Efficient Deployment of Deep Neural Networks on Multi-Accelerator Heterogeneous Edge SoCs

Deploying DNNs on System-on-Chips (SoC) with multiple heterogeneous acceleration engines is challenging, and the majority of deployment frameworks cannot fully exploit heterogeneity. We present MATCHA, a unified DNN deployment framework that generates highly concurrent schedules for parallel, heterogeneous accelerators and uses constraint programming to optimize L3/L2 memory allocation and scheduling. Using pattern matching, tiling, and mapping across individual HW units enables parallel execution and high accelerator utilization. On the MLPerf Tiny benchmark, using a SoC with two heterogeneous accelerators, MATCHA improves accelerator utilization and reduces inference latency by up to 35% with respect to the the state-of-the-art MATCH compiler.

cs.DC

CHAOS: Controlled Hardware fAult injectOr System for gem5

Fault injectors are essential tools for evaluating the reliability and resilience of computing systems. They enable the simulation of hardware and software faults to analyze system behavior under error conditions and assess its ability to operate correctly despite disruptions. Such analysis is critical for identifying vulnerabilities and improving system robustness. CHAOS is a modular, open-source, and fully configurable fault injection framework designed for the gem5 simulator. It facilitates precise and systematic fault injection across multiple architectural levels, supporting comprehensive evaluations of fault tolerance mechanisms and resilience strategies. Its high configurability and seamless integration with gem5 allow researchers to explore a wide range of fault models and complex scenarios, making CHAOS a valuable tool for advancing research in dependable and high-performance computing systems.

cs.AR

Lattice Surgery Aware Resource Analysis for the Mapping and Scheduling of Quantum Circuits for Scalable Modular Architectures

Quantum computing platforms are evolving to a point where placing high numbers of qubits into a single core comes with certain difficulties such as fidelity, crosstalk, and high power consumption of dense classical electronics. Utilizing distributed cores, each hosting logical data qubits and logical ancillas connected via classical and quantum communication channels, offers a promising alternative. However, building such a system for logical qubits requires additional optimizations, such as minimizing the amount of state transfer between cores for inter-core two-qubit gates and optimizing the routing of magic states distilled in a magic state factory. In this work, we investigate such a system and its statistics in terms of classical and quantum resources. First, we restrict our quantum gate set to a universal gate set consisting of CNOT, H, T, S, and Pauli gates. We then developed a framework that can take any quantum circuit, transpile it to our gate set using Qiskit, and then partition the qubits using the KaHIP graph partitioner to balanced partitions. Afterwards, we built an algorithm to map these graphs onto the 2D mesh of quantum cores by converting the problem into a Quadratic Assignment Problem with Fixed Assignment (QAPFA) to minimize the routing of leftover two-qubit gates between cores and the total travel of magic states from the magic state factory. Following this stage, the gates are scheduled using an algorithm that takes care of the timing of the gate set. As a final stage, our framework reports detailed statistics such as the number of classical communications, the number of EPR pairs and magic states consumed, and timing overheads for pre- and post- processing for inter-core state transfers. These results help to quantify both classical and quantum resources that are used in distributed logical quantum computing architectures.

quant-ph

Instruction-Directed MAC for Efficient Classical Communication in Scalable Multi-Chip Quantum Systems

Scalable quantum computing requires modular multi-chip architectures integrating multiple quantum cores interconnected through quantum-coherent and classical links. The classical communication subsystem is critical for coordinating distributed control operations and supporting quantum protocols such as teleportation. In this work, we consider a realization based on a wireless network-on-chip for implementing classical communication within cryogenic environments. Traditional token-based medium access control (MAC) protocols, however, incur latency penalties due to inefficient token circulation among inactive nodes. We propose the instruction-directed token MAC (ID-MAC), a protocol that leverages the deterministic nature of quantum circuit execution to predefine transmission schedules at compile time. By embedding instruction-level information into the MAC layer, ID-MAC restricts token circulation to active transmitters, thereby improving channel utilization and reducing communication latency. Simulations show that ID-MAC reduces classical communication time by up to 70% and total execution time by up to 30-70%, while also extending effective system coherence. These results highlight ID-MAC as a scalable and efficient MAC solution for future multi-chip quantum architectures.

quant-ph

Assessing the Role of Communication in Modular Multi-Core Quantum Systems

The scalability of quantum computing is constrained by the physical and architectural limitations of monolithic quantum processors. Modular multi-core quantum architectures, which interconnect multiple quantum cores (QCs) via classical and quantum-coherent links, offer a promising alternative to address these challenges. However, transitioning to a modular architecture introduces communication overhead, where classical communication plays a crucial role in executing quantum algorithms by transmitting measurement outcomes and synchronizing operations across QCs. Understanding the impact of classical communication on execution time is therefore essential for optimizing system performance. In this work, we introduce \qcomm, an open-source simulator designed to evaluate the role of classical communication in modular quantum computing architectures. \qcomm{} provides a high-level execution and timing model that captures the interplay between quantum gate execution, entanglement distribution, teleportation protocols, and classical communication latency. We conduct an extensive experimental analysis to quantify the impact of classical communication bandwidth, interconnect types, and quantum circuit mapping strategies on overall execution time. Furthermore, we assess classical communication overhead when executing real quantum benchmarks mapped onto a cryogenically-controlled multi-core quantum system. Our results show that, while classical communication is generally not the dominant contributor to execution time, its impact becomes increasingly relevant in optimized scenarios -- such as improved quantum technology, large-scale interconnects, or communication-aware circuit mappings. These findings provide useful insights for the design of scalable modular quantum architectures and highlight the importance of evaluating classical communication as a performance-limiting factor in future systems.

quant-ph

On the Impact of Classical and Quantum Communication Networks Upon Modular Quantum Computing Architecture System Performance

Modular architectures are a promising approach to scaling quantum computers beyond the limits of monolithic designs. However, non-local communications between different quantum processors might significantly impact overall system performance. In this work, we investigate the role of the network infrastructure in modular quantum computing architectures, focusing on coherence loss due to communication constraints. We analyze the impact of classical network latency on quantum teleportation and identify conditions under which it becomes a bottleneck. Additionally, we study different network topologies and assess how communication resources affect the number and parallelization of inter-core communications. Finally, we conduct a full-stack evaluation of the architecture under varying communication parameters, demonstrating how these factors influence the overall system performance. The results show that classical communication does not become a bottleneck for systems exceeding one million qubits, given current technology assumptions, even with modest clock frequencies and parallel wired interconnects. Additionally, increasing quantum communication resources generally shortens execution time, although it may introduce additional communication overhead. The optimal number of quantum links between QCores depends on both the algorithm being executed and the chosen inter-core topology. Our findings offer valuable guidance for designing modular architectures, enabling scalable quantum computing.

quant-ph

Decentralized Framework for Teleportation in Quantum Core Interconnects

Multi-core quantum computing architectures offer a promising and scalable solution to the challenges of integrating large number of qubits into existing monolithic chip design. However, the issue of transferring quantum information across the cores remains unresolved. Quantum Teleportation offers a potential approach for efficient qubit transfer, but existing methods primarily rely on centralized interconnection mechanisms for teleportation, which may limit scalability and parallel communication. We proposes a decentralized framework for teleportation in multi-core quantum computing systems, aiming to address these limitations. We introduce two variants of teleportation within the decentralized framework and evaluate their impact on reducing end-to-end communication delay and quantum circuit depth. Our findings demonstrate that the optimized teleportation strategy, termed two-way teleportation, results in a substantial 40% reduction in end-to-end communication latency for synthetic benchmarks and a 30% reduction for real benchmark applications, and 24% decrease in circuit depth compared to the baseline teleportation strategy. These results highlight the significant potential of decentralized teleportation to improve the performance of large-scale quantum systems, offering a scalable and efficient solution for future quantum architectures.

quant-ph

TeleSABRE: Layout Synthesis in Multi-Core Quantum Systems with Teleport Interconnect

Quantum circuit compilation and, in particular, efficient qubit layout synthesis is a critical challenge in modular, multi-core quantum architectures with constrained interconnects. In this work, we extend the SABRE heuristic algorithm to develop TeleSABRE, a layout synthesis approach tailored for architectures featuring teleportation-based interconnects. Unlike standard SABRE, which only introduces SWAP operations for qubit movement, TeleSABRE integrates both intracore SWAPs and teleportation-based techniques leveraging qubit teleportation and gate teleportation across cores. This enables more efficient circuit execution by reducing both inter-core communication overhead and the number of intra-core SWAPs required to allow teleportation protocols and local gate executions. Experimental results demonstrate that TeleSABRE achieves 28% reduction across various benchmarks in terms of inter-core operations while also taking into account the logistics of the teleport protocols.

quant-ph

A Survey on Deep Learning Hardware Accelerators for Heterogeneous HPC Platforms

Recent trends in deep learning (DL) have made hardware accelerators essential for various high-performance computing (HPC) applications, including image classification, computer vision, and speech recognition. This survey summarizes and classifies the most recent developments in DL accelerators, focusing on their role in meeting the performance demands of HPC applications. We explore cutting-edge approaches to DL acceleration, covering not only GPU- and TPU-based platforms but also specialized hardware such as FPGA- and ASIC-based accelerators, Neural Processing Units, open hardware RISC-V-based accelerators, and co-processors. This survey also describes accelerators leveraging emerging memory technologies and computing paradigms, including 3D-stacked Processor-In-Memory, non-volatile memories like Resistive RAM and Phase Change Memories used for in-memory computing, as well as Neuromorphic Processing Units, and Multi-Chip Module-based accelerators. Furthermore, we provide insights into emerging quantum-based accelerators and photonics. Finally, this survey categorizes the most influential architectures and technologies from recent years, offering readers a comprehensive perspective on the rapidly evolving field of deep learning acceleration.

cs.AR

A Data-Driven Approach to Dataflow-Aware Online Scheduling for Graph Neural Network Inference

Graph Neural Networks (GNNs) have shown significant promise in various domains, such as recommendation systems, bioinformatics, and network analysis. However, the irregularity of graph data poses unique challenges for efficient computation, leading to the development of specialized GNN accelerator architectures that surpass traditional CPU and GPU performance. Despite this, the structural diversity of input graphs results in varying performance across different GNN accelerators, depending on their dataflows. This variability in performance due to differing dataflows and graph properties remains largely unexplored, limiting the adaptability of GNN accelerators. To address this, we propose a data-driven framework for dataflow-aware latency prediction in GNN inference. Our approach involves training regressors to predict the latency of executing specific graphs on particular dataflows, using simulations on synthetic graphs. Experimental results indicate that our regressors can predict the optimal dataflow for a given graph with up to 91.28% accuracy and a Mean Absolute Percentage Error (MAPE) of 3.78%. Additionally, we introduce an online scheduling algorithm that uses these regressors to enhance scheduling decisions. Our experiments demonstrate that this algorithm achieves up to $3.17\times$ speedup in mean completion time and $6.26\times$ speedup in mean execution time compared to the best feasible baseline across all datasets.

cs.LG

Multi-Objective Hardware-Mapping Co-Optimisation for Multi-DNN Workloads on Chiplet-based Accelerators

The need to efficiently execute different Deep Neural Networks (DNNs) on the same computing platform, coupled with the requirement for easy scalability, makes Multi-Chip Module (MCM)-based accelerators a preferred design choice. Such an accelerator brings together heterogeneous sub-accelerators in the form of chiplets, interconnected by a Network-on-Package (NoP). This paper addresses the challenge of selecting the most suitable sub-accelerators, configuring them, determining their optimal placement in the NoP, and mapping the layers of a predetermined set of DNNs spatially and temporally. The objective is to minimise execution time and energy consumption during parallel execution while also minimising the overall cost, specifically the silicon area, of the accelerator. This paper presents MOHaM, a framework for multi-objective hardware-mapping co-optimisation for multi-DNN workloads on chiplet-based accelerators. MOHaM exploits a multi-objective evolutionary algorithm that has been specialised for the given problem by incorporating several customised genetic operators. MOHaM is evaluated against state-of-the-art Design Space Exploration (DSE) frameworks on different multi-DNN workload scenarios. The solutions discovered by MOHaM are Pareto optimal compared to those by the state-of-the-art. Specifically, MOHaM-generated accelerator designs can reduce latency by up to $96\%$ and energy by up to $96.12\%$.

cs.AR

Attention-Based Deep Reinforcement Learning for Qubit Allocation in Modular Quantum Architectures

Modular, distributed and multi-core architectures are currently considered a promising approach for scalability of quantum computing systems. The integration of multiple Quantum Processing Units necessitates classical and quantum-coherent communication, introducing challenges related to noise and quantum decoherence in quantum state transfers between cores. Optimizing communication becomes imperative, and the compilation and mapping of quantum circuits onto physical qubits must minimize state transfers while adhering to architectural constraints. The compilation process, inherently an NP-hard problem, demands extensive search times even with a small number of qubits to be solved to optimality. To address this challenge efficiently, we advocate for the utilization of heuristic mappers that can rapidly generate solutions. In this work, we propose a novel approach employing Deep Reinforcement Learning (DRL) methods to learn these heuristics for a specific multi-core architecture. Our DRL agent incorporates a Transformer encoder and Graph Neural Networks. It encodes quantum circuits using self-attention mechanisms and produce outputs through an attention-based pointer mechanism that directly signifies the probability of matching logical qubits with physical cores. This enables the selection of optimal cores for logical qubits efficiently. Experimental evaluations show that the proposed method can outperform baseline approaches in terms of reducing inter-core communications and minimizing online time-to-solution. This research contributes to the advancement of scalable quantum computing systems by introducing a novel learning-based heuristic approach for efficient quantum circuit compilation and mapping.

quant-ph

Assessing the Role of Communication in Scalable Multi-Core Quantum Architectures

Multi-core quantum architectures offer a solution to the scalability limitations of traditional monolithic designs. However, dividing the system into multiple chips introduces a critical bottleneck: communication between cores. This paper introduces qcomm, a simulation tool designed to assess the impact of communication on the performance of scalable multi-core quantum architectures. Qcomm allows users to adjust various architectural and physical parameters of the system, and outputs various communication metrics. We use qcomm to perform a preliminary study on how these parameters affect communication performance in a multi-core quantum system.

quant-ph

Deep Reinforcement Learning based Online Scheduling Policy for Deep Neural Network Multi-Tenant Multi-Accelerator Systems

Currently, there is a growing trend of outsourcing the execution of DNNs to cloud services. For service providers, managing multi-tenancy and ensuring high-quality service delivery, particularly in meeting stringent execution time constraints, assumes paramount importance, all while endeavoring to maintain cost-effectiveness. In this context, the utilization of heterogeneous multi-accelerator systems becomes increasingly relevant. This paper presents RELMAS, a low-overhead deep reinforcement learning algorithm designed for the online scheduling of DNNs in multi-tenant environments, taking into account the dataflow heterogeneity of accelerators and memory bandwidths contentions. By doing so, service providers can employ the most efficient scheduling policy for user requests, optimizing Service-Level-Agreement (SLA) satisfaction rates and enhancing hardware utilization. The application of RELMAS to a heterogeneous multi-accelerator system composed of various instances of Simba and Eyeriss sub-accelerators resulted in up to a 173% improvement in SLA satisfaction rate compared to state-of-the-art scheduling techniques across different workload scenarios, with less than a 1.5% energy overhead.

cs.AR

Towards Fair and Firm Real-Time Scheduling in DNN Multi-Tenant Multi-Accelerator Systems via Reinforcement Learning

This paper addresses the critical challenge of managing Quality of Service (QoS) in cloud services, focusing on the nuances of individual tenant expectations and varying Service Level Indicators (SLIs). It introduces a novel approach utilizing Deep Reinforcement Learning for tenant-specific QoS management in multi-tenant, multi-accelerator cloud environments. The chosen SLI, deadline hit rate, allows clients to tailor QoS for each service request. A novel online scheduling algorithm for Deep Neural Networks in multi-accelerator systems is proposed, with a focus on guaranteeing tenant-wise, model-specific QoS levels while considering real-time constraints.

cs.AR