SearcharxivSearch

arXiv subjects

Chaojie Zhang

Publications and source records attributed to Chaojie Zhang.

At least 19 recordsLinked to original sources

Slasher: Power Flexibility for Cloud Datacenters

Datacenters consume many megawatts of power, and regularly encounter scenarios that require modulating their power draw. These scenarios include datacenter infrastructure failures, power grid failures, grid services, and more, spanning a diverse range of requirements in terms of the power magnitude, the scope of the reduction, the notice time, and other dimensions. To address these scenarios, we have built Slasher, a general system for modulating the power of \azure datacenters to handle scenarios ranging from individual racks to regional multi-datacenter grid events. Slasher coordinates datacenter resources with the goal of meeting power targets while minimizing negative impact on hosted workloads. In this paper, we review the main power modulation scenarios, characterize the power reduction levers using data from production cloud datacenters, describe Slasher's system architecture, and formulate the cloud datacenter power modulation control problem. We also develop a high-fidelity datacenter simulator and propose a workload impact model, using them to design and evaluate power control algorithms.

cs.DC

Architectural Implications of Agentic AI Workflows

Agentic AI is emerging in datacenters, but its architectural implications remain unexplored. We organize agentic workflows in a taxonomy and present its first architectural characterization with a production study at Microsoft Azure and a controlled study of open-source frameworks. We show that agentic execution is fragmented and heterogeneous. Requests expand into a workflow of LLM inferences, tool invocations, and orchestration decisions that repeatedly cross the CPU-GPU boundary. Our taxonomy explains how this fragmentation turns into resource demand. As orchestration and tools run on the host, the CPU sits on the critical path. Execution structure sets the load over time, which stays low with sudden spikes. Model composition sets how evenly the workflow uses the GPUs. Diversity in tasks and tools widens this range even further. These characteristics expose architectural mismatches of conventional uniform servers. Fragmented execution strands CPU and GPU capacity despite bursty demand. Different software roles make homogeneous CPU provisioning inefficient. Finally, multiplexing many agents onto shared cores degrades microarchitectural locality. Guided by our findings, we derive implications for agentic servers and examine them through Agora, our prototype for commodity servers. Agora dynamically harvests idle CPU cores for co-located throughput work, while protecting agentic tail latency against tool spikes. It oversubscribes GPU memory by placing more agents on each GPU, prefetching the next agent's state to hide swap latency. To match the machine to the heterogeneous roles, Agora pools cores by role and applies affinity-aware scheduling to restore locality. It automatically tunes mechanisms to the workload. Agora improves utilization and server throughput while preserving agent tail latency. Our insights also identify key directions for future server architectures for agentic AI.

cs.AI

AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane Policies

The efficiency of a datacenter rests on its control plane policies. Designing these policies is increasingly hard: the hardware-software stack grows fast, the design space is vast and interdependent, and prototyping a single policy takes months. Agentic AI promises to automate this search. Off the shelf, however, it falls short on three fronts. It is not formal: with no structured, searchable statement of the problem, the search has little structure to exploit and hard constraints are not guaranteed. It is not transferable: each task is solved from scratch, so nothing learned on one task carries to the next. Finally, it is not systematic: relying on the LLM as the sole source of candidates, it explores a narrow slice of the design space and settles into local optima. We introduce AtumAI, a framework that generates datacenter control-plane policies with agentic AI, making the process formal, transferable, and systematic. From a goal stated in plain language, AtumAI autonomously proposes, tests, and refines candidate policies until one satisfies the request. It does so through two components. The Datacenter Task Compiler automates problem formulation: it compiles the request into a formal, machine-checkable, and searchable specification of the task's objectives, constraints, decision variables, and evaluation methodology. The Evolutionary Design Discovery Loop then searches this specification, expanding the search beyond the LLM itself via a diffusion model, an evolutionary algorithm, and a surrogate model. Together, they reduce onboarding a new task from months of engineering to writing its description. We evaluate AtumAI on three control-plane tasks with distinct problem scopes, design spaces, and trade-offs: workload placement, resource scaling, and power management. Across all tasks, the policies generated by AtumAI consistently outperform expert-engineered baselines.

cs.AI

CWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy Farms

AI power demand is growing at an unprecedented rate while power grids are often ailing and struggle to keep up. Grid expansion comes with high capital expenditure and long-distance transmission losses, yet there is abundant renewable energy at the source, just not matched to demand. This paper proposes a complementary AI infrastructure deployment model, AI Greeninferencing, that brings modular AI compute to renewable energy sources, focusing on wind, allowing AI footprint expansion, generating local behind-the-meter demand for renewable sites, and helping ease the growing strain on power utilities. Our feasibility analysis shows that 890+ GW of wind capacity lies within 50 ms network round trip time of Azure data centers, and that site-wise right-sizing combined with spatial complementarity of wind energy keeps aggregate fleet utilization on par with traditional deployments. To serve inference requests under variable wind power, we build CWind, a lightweight, reactive, and workload-agnostic AI inference router that uses only real-time signals: inference latency, KV-cache utilization, and queue depth, to dynamically configure sites and distribute requests. Evaluated on a real 64-GPU A100 testbed emulating three wind-powered sites with Azure production traces, CWind reduces P99 end-to-end latency by up to 52% over the strongest contender (also our idea) and by up to 98% over baselines such as power-capping and GPU idling, with consistent gains across workload types, load levels, and GPU generations.

cs.DC

Tunable Nonlocal $ZZ$ Interaction for Remote Controlled-Z Gates Between Distributed Fixed-Frequency Qubits

Scaling superconducting quantum processors toward fault-tolerant operation will likely require architectures that extend beyond monolithic chips. Modular processors connected by low-loss superconducting links provide a promising route, but implementing entangling gates between remote fixed-frequency qubits remains challenging. Here we propose a distributed architecture in which two synchronously controlled double-transmon couplers mediate the interaction between fixed-frequency transmons in separate packages connected by a 25-cm coaxial cable. The scheme activates a tunable nonlocal $ZZ$ interaction on demand while suppressing residual static coupling, allowing the superconducting link to function as a gate-native interconnect rather than solely as a state-transfer channel. Circuit-level simulations show an on/off ratio exceeding $10^6$ and a remote controlled-Z gate with a projected coherent fidelity of $99.99\%$ under experimentally relevant parameters. Open-system simulations further indicate that, within the representative Markovian noise model considered here, endpoint-qubit decoherence is the largest contribution to gate infidelity, while photon loss in the retained cable modes remains smaller but non-negligible. These results identify DTC-mediated tunable nonlocal coupling as a promising gate primitive for modular superconducting processors based on fixed-frequency qubits.

quant-ph

Multi-GeV Electron Combs from a Plasma Wakefield Accelerator

Plasma accelerators now produce GeV-class electron beams with brightness and stability sufficient to drive free-electron lasers. Beyond this, they possess a unique yet largely unexplored capability: shaping the phase space of the beam in situ during injection, on femtosecond or shorter timescales. Here we demonstrate this capability by generating a multi-GeV electron comb comprising more than ten microbunches simultaneously separated in both energy and time. Periodic pinching of the drive beam inside its self-excited plasma wake sequentially injects microbunches via ionization of embedded helium atoms at successive betatron oscillations, while the gently varying plasma density maps each bunchlet to a distinct wake phase, compressing electrons trapped over a ~17 cm region into a comb only micrometers long. Individual microbunches exhibit percent-level energy spreads, energy spacing up to ten percent, and contain several picocoulomb charge. The percent-level spreads and parabolic energy-spacing trend provide experimental evidence for sub-femtosecond microbunch durations and few-femtosecond separations as revealed by beam-loading analysis and confirmed by particle-in-cell simulations. This work demonstrates femtosecond, in-situ phase-space shaping in plasma accelerators, paving the way for electron beams with tailored energy-time structure.

physics.acc-ph

TeV Electron Beams from Plasma Acceleration via Regenerative Cascading

Plasma accelerators sustain gradients orders of magnitude higher than conventional radiofrequency machines, but most proposed paths to TeV energies still require tens of stages, each demanding sub-micrometer alignment, femtosecond synchronization, and precise matching of the accelerating trailing bunch. Here we introduce plasma wakefield acceleration via regenerative cascading, in which each stage self-injects a fresh trailing electron bunch and the accelerated trailing bunch becomes the driver for the next stage. This approach has several advantages: energy multiplication instead of addition; automatic alignment, synchronization, and matching of the trailing bunch to the wake; and trailing bunch brightness reset in each stage. Particle-in-cell simulations show the generation of a 1.1 TeV electron beam with ~0.3% rms energy spread and 0.12 nC charge from a two-stage, sub-kilometer plasma accelerator driven by a 45 GeV, 100 nC beam. The low energy spread is achieved via dynamic beam loading in the evolving wake of the post-depletion driver that acts as a built-in energy dechirper.

physics.acc-ph

Designing Datacenter Power Delivery Hierarchies for the AI Era

Demand for AI accelerators is rapidly increasing rack power density, with projections approaching 1MW per deployment by 2027. This poses a major challenge for datacenter power delivery designers. As power densities increase, a datacenter designed for a different target density may strand power, i.e., may be unable to use all the power that its delivery hierarchy has provisioned. Designs must remain efficient over long datacenter lifetimes and multiple hardware generations. Power utilization is particularly important as grid power capacity is a scarce resource in the AI era. Designing an efficient power delivery hierarchy for the long run is difficult because rack placement feasibility, workload impact, and cost depend jointly on electrical topology, deployment granularity, placement policy, power oversubscription, and workload mix. Moreover, each of these factors evolve over time, have inter-dependencies across multiple resource dimensions, and generally do not lend themselves to closed-form analysis. To address this challenge, we develop a framework for evaluating datacenter power delivery designs using throughput, power, and cost metrics over realistic arrival, oversubscription, and decommissioning sequences. The framework combines projection models for GPU, compute, and storage deployments with operational factors grounded in production data from Microsoft Azure. Our results show that multi-resource stranding materially changes deployable capacity, effective capital expenditure, and delivered performance, and quantify how rising density from rack- and pod-scale AI systems shapes these outcomes. For AI datacenter design, the relevant planning objective is not installed megawatts, but deployable capacity over time.

cs.DC

Strong-field focusing of high-energy particles in beam-multifoil collisions

Extreme beams of charged particles and photons, reaching ultrahigh densities or producing intense gamma-ray bursts, are central to accelerator physics, laboratory astrophysics, and strong-field quantum electrodynamics research. Yet their generation is hindered by conventional focusing methods at multi-GeV energies that rely on massive magnetic assemblies, limiting compactness and attainable density. Here we report the first experimental observation of a fundamentally new focusing mechanism, in which a high-energy charged-particle beam is focused by its own magnetic field reflected from a stack of thin metallic foils via near-field coherent-transition-radiation. The experiment, performed at SLAC's FACET-II facility, reveals strong, cumulative focusing across a broad range of beam configurations, enabled by the delivered 10 GeV, 1 nC, 10 Hz electron beam. The measurements closely agree with predictions from an analytical model and particle-in-cell simulations. These results demonstrate that multifoil focusing is a remarkably straightforward, self-aligned approach to the generation of ultrahigh density beams, opening a path to explore unprecedented regimes of beam-matter interaction and high-energy radiation.

physics.acc-ph

Fast CZ Gate via Energy-Level Engineering in Superconducting Qubits with a Tunable Coupler

In superconducting quantum circuits, decoherence errors in qubits constitute a critical factor limiting quantum gate performance. To mitigate decoherence-induced gate infidelity, rapid implementation of quantum gates is essential. Here we propose a scheme for rapid controlled-Z (CZ) gate implementation through energy-level engineering, which leverages Rabi oscillations between the $\left|11\right\rangle$ state and the non-computational state in a tunable-coupler architecture. Numerical simulations achieved a $\mathrm{22~ns}$ nonadiabatic CZ gate with fidelity over $99.99\%$. We further investigated the performance of the CZ gate in the presence of anharmonicity offsets. The results demonstrate that a high-fidelity CZ gate with an error rate below $10^{-4}$ remains achievable even with finite anharmonicity variations. Furthermore, the detrimental impact of spectator qubits in different quantum states on the fidelity of CZ gate is effectively suppressed by incorporating a tunable coupler. This scheme exhibits potential for extending the circuit execution depth constrained by coherence time limitations.

quant-ph

StreamWise: Serving Multi-Modal Generation in Real-Time at Scale

Advances in multi-modal generative models are enabling new applications, from storytelling to automated media synthesis. Most current workloads generate simple outputs (e.g., image generation from a prompt) in batch mode, often requiring several seconds even for basic results. Serving real-time multi-modal workflows at scale is costly and complex, requiring efficient coordination of diverse models (each with unique resource needs) across language, audio, image, and video, all under strict latency and resource constraints. We tackle these challenges through the lens of real-time podcast video generation, integrating LLMs, text-to-speech, and video-audio generation. To meet tight SLOs, we design an adaptive, modular serving system, StreamWise, that dynamically manages quality (e.g., resolution, sharpness), model/content parallelism, and resource-aware scheduling. We leverage heterogeneous hardware to maximize responsiveness and efficiency. For example, the system can lower video resolution and allocate more resources to early scenes. We quantify the trade-offs between latency, cost, and quality. The cheapest setup generates a 10-minute podcast video on A100 GPUs in 1.4 hours (8.4x slower than the real-time) for less than \$25. StreamWise enables high-quality real-time streaming with a sub-second startup delay under $45.

cs.DC

No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha

Deploying million-token Large Language Models (LLMs) is challenging because production workloads are highly heterogeneous, mixing short queries and long documents. This heterogeneity, combined with the quadratic complexity of attention, creates severe convoy effects where long-running requests stall short, interactive ones, degrading system responsiveness. We present Medha, a serving system that eliminates these convoys by introducing fine-grained, preemptive scheduling to LLM inference. Medha makes preemption practical with a co-designed set of mechanisms -- including Adaptive Chunking and Stream Pipeline Parallel that overcome the perceived inefficiencies and scaling challenges of chunking. Additionally, we present a new parallelism strategy KV-Cache Parallelism to reduce the decode latency and afford interactivity despite very long context. These mechanisms are orchestrated by a Length-Aware Relative Slack (LARS) scheduler, a deadline and heterogeneity-aware scheduling policy that prevents both the convoy effect and the starvation that plagues simpler policies. Under a heterogeneous workload, Medha improves throughput by 5.7x while reducing median and 99th percentile latency by 30x and 174x, respectively, compared to state-of-the-art non-preemptive systems.

cs.LG

Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework

The rapid rise of large language models (LLMs) has been driving an enormous demand for AI inference infrastructure, mainly powered by high-end GPUs. While these accelerators offer immense computational power, they incur high capital and operational costs due to frequent upgrades, dense power consumption, and cooling demands, making total cost of ownership (TCO) for AI datacenters a critical concern for cloud providers. Unfortunately, traditional datacenter lifecycle management (designed for general-purpose workloads) struggles to keep pace with AI's fast-evolving models, rising resource needs, and diverse hardware profiles. In this paper, we rethink the AI datacenter lifecycle scheme across three stages: building, hardware refresh, and operation. We show how design choices in power, cooling, and networking provisioning impact long-term TCO. We also explore refresh strategies aligned with hardware trends. Finally, we use operation software optimizations to reduce cost. While these optimizations at each stage yield benefits, unlocking the full potential requires rethinking the entire lifecycle. Thus, we present a holistic lifecycle management framework that coordinates and co-optimizes decisions across all three stages, accounting for workload dynamics, hardware evolution, and system aging. Our system reduces the TCO by up to 40\% over traditional approaches. Using our framework we provide guidelines on how to manage AI datacenter lifecycle for the future.

cs.AI

AI Greenferencing: Routing AI Inferencing to Green Modular Data Centers with Heron

AI power demand is growing unprecedentedly thanks to the high power density of AI compute and the emerging inferencing workload. On the supply side, abundant wind power is waiting for grid access in interconnection queues. In this light, this paper argues bringing AI workload to modular compute clusters co-located in wind farms. Our deployment right-sizing strategy makes it economically viable to deploy more than 6 million high-end GPUs today that could consume cheap, green power at its source. We built Heron, a cross-site software router, that could efficiently leverage the complementarity of power generation across wind farms by routing AI inferencing workload around power drops. Using 1-week ofcoding and conversation production traces from Azure and (real) variable wind power traces, we show how Heron improves aggregate goodput of AI compute by up to 80% compared to the state-of-the-art.

cs.DC

EDA-Q: Electronic Design Automation for Superconducting Quantum Chip

Electronic Design Automation (EDA) plays a crucial role in classical chip design and significantly influences the development of quantum chip design. However, traditional EDA tools cannot be directly applied to quantum chip design due to vast differences compared to the classical realm. Several EDA products tailored for quantum chip design currently exist, yet they only cover partial stages of the quantum chip design process instead of offering a fully comprehensive solution. Additionally, they often encounter issues such as limited automation, steep learning curves, challenges in integrating with actual fabrication processes, and difficulties in expanding functionality. To address these issues, we developed a full-stack EDA tool specifically for quantum chip design, called EDA-Q. The design workflow incorporates functionalities present in existing quantum EDA tools while supplementing critical design stages such as device mapping and fabrication process mapping, which users expect. EDA-Q utilizes a unique architecture to achieve exceptional scalability and flexibility. The integrated design mode guarantees algorithm compatibility with different chip components, while employing a specialized interactive processing mode to offer users a straightforward and adaptable command interface. Application examples demonstrate that EDA-Q significantly reduces chip design cycles, enhances automation levels, and decreases the time required for manual intervention. Multiple rounds of testing on the designed chip have validated the effectiveness of EDA-Q in practical applications.

cs.ET

Design Initiative for a 10 TeV pCM Wakefield Collider

This document outlines a community-driven Design Study for a 10 TeV pCM Wakefield Accelerator Collider. The 2020 ESPP Report emphasized the need for Advanced Accelerator R\&D, and the 2023 P5 Report calls for the ``delivery of an end-to-end design concept, including cost scales, with self-consistent parameters throughout." This Design Study leverages recent experimental and theoretical progress resulting from a global R\&D program in order to deliver a unified, 10 TeV Wakefield Collider concept. Wakefield Accelerators provide ultra-high accelerating gradients which enables an upgrade path that will extend the reach of Linear Colliders beyond the electroweak scale. Here, we describe the organization of the Design Study including timeline and deliverables, and we detail the requirements and challenges on the path to a 10 TeV Wakefield Collider.

physics.acc-ph

TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms

The rising demand for generative large language models (LLMs) poses challenges for thermal and power management in cloud datacenters. Traditional techniques often are inadequate for LLM inference due to the fine-grained, millisecond-scale execution phases, each with distinct performance, thermal, and power profiles. Additionally, LLM inference workloads are sensitive to various configuration parameters (e.g., model parallelism, size, and quantization) that involve trade-offs between performance, temperature, power, and output quality. Moreover, clouds often co-locate SaaS and IaaS workloads, each with different levels of visibility and flexibility. We propose TAPAS, a thermal- and power-aware framework designed for LLM inference clusters in the cloud. TAPAS enhances cooling and power oversubscription capabilities, reducing the total cost of ownership (TCO) while effectively handling emergencies (e.g., cooling and power failures). The system leverages historical temperature and power data, along with the adaptability of SaaS workloads, to: (1) efficiently place new GPU workload VMs within cooling and power constraints, (2) route LLM inference requests across SaaS VMs, and (3) reconfigure SaaS VMs to manage load spikes and emergency situations. Our evaluation on a large GPU cluster demonstrates significant reductions in thermal and power throttling events, boosting system efficiency.

cs.DC

DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency

The rapid evolution and widespread adoption of generative large language models (LLMs) have made them a pivotal workload in various applications. Today, LLM inference clusters receive a large number of queries with strict Service Level Objectives (SLOs). To achieve the desired performance, these models execute on power-hungry GPUs causing the inference clusters to consume large amount of energy and, consequently, result in excessive carbon emissions. Fortunately, we find that there is a great opportunity to exploit the heterogeneity in inference compute properties and fluctuations in inference workloads, to significantly improve energy-efficiency. However, such a diverse and dynamic environment creates a large search-space where different system configurations (e.g., number of instances, model parallelism, and GPU frequency) translate into different energy-performance trade-offs. To address these challenges, we propose DynamoLLM, the first energy-management framework for LLM inference environments. DynamoLLM automatically and dynamically reconfigures the inference cluster to optimize for energy and cost of LLM serving under the service's performance SLOs. We show that at a service-level, DynamoLLM conserves 53% energy and 38% operational carbon emissions, and reduces 61% cost to the customer, while meeting the latency SLOs.

cs.AI