SearcharxivSearch

arXiv subjects

Zongwei Zhu

Publications and source records attributed to Zongwei Zhu.

4 recordsLinked to original sources

Beacon: LLM Multi-Agent Driven Hardware Design Space Exploration for Heterogeneous Multi-Chiplet Deep Learning Accelerators

Heterogeneous multi-chiplet accelerators allow chiplets to be configured independently to better match different operator characteristics and improve inference efficiency. However, heterogeneity makes simulator evaluation expensive, limiting the number of iterations affordable for hardware design space exploration (HW-DSE). Mainstream data-driven methods rely mainly on final metrics and a few predefined states, and require many search iterations to implicitly learn the relationships between input parameters and optimization objectives, making them less effective in this setting. In practice, evaluators also generate detailed reports on execution timelines, resource utilization, memory accesses, and communication behavior. Large language models (LLMs) can combine domain knowledge with these reports to explicitly identify bottleneck locations, degradation causes, and parameter adjustment directions, thereby improving each design decision under limited iteration budgets. Based on this observation, we propose Beacon, a report-driven LLM multi-agent framework for heterogeneous multi-chiplet HW-DSE. Beacon employs hierarchical agents for bottleneck localization, root-cause diagnosis, and hardware candidate generation, together with an Analysis Toolbox and RAG memory for closed-loop search. Under the same limited iteration budget, Beacon reduces the composite latency-energy-monetary-cost objective by 25.1\%--93.5\% compared with random search, Bayesian optimization, and reinforcement learning.

cs.AR

Compass: Co-Exploration of Mapping and Hardware for Heterogeneous Multi-Chiplet Accelerators Targeting LLM Inference Service Workloads

Large language models (LLMs) bring huge computational demands, which makes multi-chiplet accelerators that can integrate large-scale computing resources a powerful solution. However, existing design space exploration (DSE) efforts for such accelerators primarily focus on traditional CNN/Transformer workloads and fall short in supporting the highly dynamic behavior of real-world LLM inference services. This dynamic nature manifests in two key aspects: 1) Mixed request types: the prefill and decode phases exhibit significantly different computational patterns and are frequently interleaved by modern system-level service schedulers; 2) Variable sequence lengths: the sequence length differences across requests can span several orders of magnitude, rendering padding-based assumptions inefficient. Moreover, many prior works assume homogeneous chiplets and overlook the potential beneficial interaction between LLM dynamics and heterogeneous chiplet architectures. To bridge this gap, we introduce Compass, a co-exploration framework designed to optimize mapping strategies and hardware design for multi-chiplet accelerators, specifically tailored for dynamic LLM workloads. First, we propose a computation execution graph-based mapping encoding scheme that decouples micro-batch and layer dimensions, enabling fine-grained execution control on heterogeneous chiplets and flexibly representing various parallelism strategies. Second, based on this scheme, we develop the Compass framework itself, which integrates an evaluation engine, a mapping generation engine based on genetic algorithm, and a hardware sampling engine based on Bayesian optimization, enabling fast and flexible cross-level co-design. Compared with the SOTA DSE works Gemini and MOHaM, Compass reduces latency by 63.92\% and energy by 40.32\% on average in various scenarios, with only a 3.11\% increase in monetary cost.

cs.AR

DeltaFS: Pursuing Zero Update Overhead via Metadata-Enabled Delta Compression for Log-structured File System on Mobile Devices

Data compression has been widely adopted to release mobile devices from intensive write pressure. Delta compression is particularly promising for its high compression efficacy over conventional compression methods. However, this method suffers from non-trivial system overheads incurred by delta maintenance and read penalty, which prevents its applicability on mobile devices. To this end, this paper proposes DeltaFS, a metadata-enabled Delta compression on log-structured File System for mobile devices, to achieve utmost compressing efficiency and zero hardware costs. DeltaFS smartly exploits the out-of-place updating ability of Log-structured File System (LFS) to alleviate the problems of write amplification, which is the key bottleneck for delta compression implementation. Specifically, DeltaFS utilizes the inline area in file inodes for delta maintenance with zero hardware cost, and integrates an inline area manage strategy to improve the utilization of constrained inline area. Moreover, a complimentary delta maintenance strategy is incorporated, which selectively maintains delta chunks in the main data area to break through the limitation of constrained inline area. Experimental results show that DeltaFS substantially reduces write traffics by up to 64.8\%, and improves the I/O performance by up to 37.3\%.

cs.DC

HADFL: Heterogeneity-aware Decentralized Federated Learning Framework

Federated learning (FL) supports training models on geographically distributed devices. However, traditional FL systems adopt a centralized synchronous strategy, putting high communication pressure and model generalization challenge. Existing optimizations on FL either fail to speedup training on heterogeneous devices or suffer from poor communication efficiency. In this paper, we propose HADFL, a framework that supports decentralized asynchronous training on heterogeneous devices. The devices train model locally with heterogeneity-aware local steps using local data. In each aggregation cycle, they are selected based on probability to perform model synchronization and aggregation. Compared with the traditional FL system, HADFL can relieve the central server's communication pressure, efficiently utilize heterogeneous computing power, and can achieve a maximum speedup of 3.15x than decentralized-FedAvg and 4.68x than Pytorch distributed training scheme, respectively, with almost no loss of convergence accuracy.

cs.LG