SearcharxivSearch

arXiv subjects

Boris Vaisband

Publications and source records attributed to Boris Vaisband.

2 recordsLinked to original sources

SHIFT: Dynamic Compute Relocation Framework for Communication-Aware Chiplet-Based Systems

The increasing communication complexity of large-scale heterogeneous systems has motivated runtime methodologies for communication-aware workload placement and routing optimization. These communication limitations are addressed in this paper by proposing SHIFT, a novel topology-agnostic approach that transfers compute node context and data to a more suitably positioned node, rather than only shifting data as in conventional networks-on-chip. The proposed strategy is evaluated on a chiplet-based architecture utilizing a fine-pitch integration platform featuring multiple bandwidth-domains for heterogeneous workloads. The proposed architecture employs multi-layered routing between functional or memory chiplets and utility chiplets, which serve as intelligent nodes for routing and compute relocation. Adaptive scheduling and routing utilize a modified shortest-path algorithm for large-scale systems, complemented by a lightweight ML-assisted policy that infers traffic conditions to improve adaptivity. To establish a performance baseline, the initial assessment uses random instruction vectors and data patterns to evaluate the fundamental capabilities of SHIFT. Simulation results exhibit successful relocations over total trials ranging from 75.2% to 97.9% across configurations, with average latency improvements of 16.4%-62.5% and a maximum of 76.8%. In addition, throughput is improved by up to 12.5x, power dissipation per unit area is reduced by ~8%, energy-per-bit is reduced by up to 58.3%, and performance is improved by 18%. To evaluate efficiency under high logic and data density, the framework was tested on standard LLM workloads. Results exhibit average improvements of 4.9x, 5.9x, and 1.8x in runtime, throughput, and energy-efficiency, respectively, surpassing state-of-the-art wafer-scale LLM services and demonstrating compatibility with large-scale platforms and applications.

cs.AR

Power Delivery for Ultra-Large-Scale Applications on Si-IF

In recent years, with the rise of artificial intelligence and big data, there is an even greater demand for scaling out computing and memory capacity. Silicon interconnect fabric (Si-IF), a wafer-scale integration platform, promotes a paradigm shift in packaging features and enables ultra-large-scale systems, while significantly improving communication bandwidth and latency. Such systems are expected to dissipate tens of kilowatts of power. Designing an efficient and robust power delivery methodology for these high power applications is a key challenge in the enablement of the Si-IF platform. Based on several figure-of-merit parameters, an efficient power delivery methodology is matched with each of three candidate applications on the Si-IF, namely, artificial intelligence accelerators, high-performance computing, and neuromorphic computing. The proposed power delivery approaches were simulated and exhibit compatibility with the relevant ultra-large-scale application on Si-IF. The simulation results confirm that the dedicated power delivery topologies can support ultra-large-scale applications on the SI-IF.

cs.AR