SearcharxivSearch

arXiv subjects

Andrew A. Chien

Publications and source records attributed to Andrew A. Chien.

13 recordsLinked to original sources

DCGen 1.1 Technical Report: Generating Datacenter Configurations (including IT, Power, Cooling)

Diversification of digital applications and workloads has driven the development of diverse datacenter architectures on ever-larger scales. These datacenters consist of complex IT, power, and cooling systems with interdependencies that influence configuration and performance. As datacenters scale and power density increase, designing realistic models becomes more difficult, particularly for research, because it requires understanding all layers of the datacenter and how they interact. Consequently, many studies rely on outdated or unrealistic designs. To support research in datacenter hardware design principles, operational dynamics, cooling mechanisms, and interactions of these facilities with the electrical grid, we have designed DCGen, a tool which can generate a variety of datacenter configurations (including IT hardware, cooling and power distribution infrastructures) at various electrical power, compute capability, and area targets.The tool captures power and space characteristics of IT, cooling, and power infrastructures at both the rack and datacenter levels, enabling modeling of power, energy, and space. DCGen leverages specific use cases such as AI training, AI inference, and cloud services, to select reference and canonical IT hardware configurations, producing realistic mixes of server types. It can target datacenter scale in terms of both power (e.g., 10 MW, 100 MW, 1 GW) and compute capability. For cooling and power distribution infrastructures, DCGen chooses components from a production equipment catalog that optimizes for space or power efficiency while meeting the datacenter capacity requirements. This tool supports research using realistic datacenter designs through ``what-if'' scenario exploration, including studies of power density evolution over time, grid interconnection capacity planning, datacenter-grid interactions, and space management.

cs.DC

Distribution and Management of Datacenter Load Decoupling

The exploding power consumption of AI and cloud datacenters (DCs) intensifies the long-standing concerns about their carbon footprint, especially because DCs' need for constant power clashes with volatile renewable generation needed for grid decarbonization. DC flexibility (a.k.a. load adaptation) is a key to reducing DC carbon emissions by improving grid renewable absorption. DC flexibility can be created, without disturbing datacenter capacity by decoupling a datacenter's power capacity and grid load with a collection of energy resources. Because decoupling can be costly, we study how to best distribute and manage decoupling to maximize benefits for all. Key considerations include site variation and datacenter-grid cooperation. We first define and compute the power and energy needs of datacenter load decoupling, and then we evaluate designed distribution and management approaches. Evaluation shows that optimized distribution can deliver >98% of the potential grid carbon reduction with 70% of the total decoupling need. For management, DC-grid cooperation (2-way sharing and control vs. 1-way info sharing) enables 1.4x grid carbon reduction. Finally, we show that decoupling may be economically viable, as on average datacenters can get power cost and carbon emissions benefits greater than their local costs of decoupling. However, skew across sites suggests grid intervention may be required.

cs.DC

Modeling the Carbon Footprint of HPC: The Top 500 and EasyC

Climate change is a critical concern for HPC systems, but GHG protocol carbon-emission accounting methodologies are difficult for a single system, and effectively infeasible for a collection of systems. As a result, there is no HPC-wide carbon reporting, and even the largest HPC sites do not do GHG protocol reporting. We assess the carbon footprint of HPC, focusing on the Top 500 systems. The key challenge lies in modeling the carbon footprint with limited data availability. With the disclosed top500 website data, and using a new tool, EasyC, we were able to model the operational carbon of 391 HPC systems and the embodied carbon of 283 HPC systems. We further show how this coverage can be enhanced by exploiting additional public information. With improved coverage, then interpolation is used to produce the first carbon footprint estimates of the Top 500 HPC systems. They are 1.4 million MT CO2e operational carbon (1 Year) and 1.9 million MT CO2e embodied carbon. We also project how the Top 500's carbon footprint will increase through 2030. A key enabler is the EasyC tool which models carbon footprint with only a few data metrics. We explore availability of data and enhancement, showing that coverage can be increased to 98% of Top 500 systems for operational and 80.8% of the systems for embodied emissions.

cs.DC

How Fast Can Graph Computations Go on Fine-grained Parallel Architectures

Large-scale graph problems are of critical and growing importance and historically parallel architectures have provided little support. In the spirit of co-design, we explore the question, How fast can graph computing go on a fine-grained architecture? We explore the possibilities of an architecture optimized for fine-grained parallelism, natural programming, and the irregularity and skew found in real-world graphs. Using two graph benchmarks, PageRank (PR) and Breadth-First Search (BFS), we evaluate a Fine-Grained Graph architecture, UpDown, to explore what performance codesign can achieve. To demonstrate programmability, we wrote five variants of these algorithms. Simulations of up to 256 nodes (524,288 lanes) and projections to 16,384 nodes (33M lanes) show the UpDown system can achieve 637K GTEPS PR and 989K GTEPS BFS on RMAT, exceeding the best prior results by 5x and 100x respectively.

cs.DC

UpDown: Programmable fine-grained Events for Scalable Performance on Irregular Applications

Applications with irregular data structures, data-dependent control flows and fine-grained data transfers (e.g., real-world graph computations) perform poorly on cache-based systems. We propose the UpDown accelerator that supports fine-grained execution with novel architecture mechanisms - lightweight threading, event-driven scheduling, efficient ultra-short threads, and split-transaction DRAM access with software-controlled synchronization. These hardware primitives support software programmable events, enabling high performance on diverse data structures and algorithms. UpDown also supports scalable performance; hardware replication enables programs to scale up performance. Evaluation results show UpDown's flexibility and scalability enable it to outperform CPUs on graph mining and analytics computations by up to 116-195x geomean speedup and more than 4x speedup over prior accelerators. We show that UpDown generates high memory parallelism (~4.6x over CPU) required for memory intensive graph computations. We present measurements that attribute the performance of UpDown (23x architectural advantage) to its individual architectural mechanisms. Finally, we also analyze the area and power cost of UpDown's mechanisms for software programmability.

cs.AR

Modeling Performance of Data Collection Systems for High-Energy Physics

Exponential increases in scientific experimental data are outstripping the rate of progress in silicon technology. As a result, heterogeneous combinations of architectures and process or device technologies are increasingly important to meet the computing demands of future scientific experiments. However, the complexity of heterogeneous computing systems requires systematic modeling to understand performance. We present a model which addresses this need by framing key aspects of data collection pipelines and constraints, and combines them with the important vectors of technology that shape alternatives, computing metrics that allow complex alternatives to be compared. For instance, a data collection pipeline may be characterized by parameters such as sensor sampling rates, amount of data collected, and the overall relevancy of retrieved samples. Alternatives to this pipeline are enabled by hardware development vectors including advancing CMOS, GPUs, neuromorphic computing, and edge computing. By calculating metrics for each alternative such as overall F1 score, power, hardware cost, and energy expended per relevant sample, this model allows alternate data collection systems to be rigorously compared. To demonstrate this model's capability, we apply it to the CMS experiment (and planned HL-LHC upgrade) to evaluate and compare the application of novel technologies in the data acquisition system (DAQ). We demonstrate that improvements to early stages in the DAQ are highly beneficial, greatly reducing the resources required at later stages of processing (such as a 60% power reduction) and increasing the amount of relevant data retrieved from the experiment per unit power (improving from 0.065 to 0.31 samples/kJ) However, we predict further advances will be required in order to meet overall power and cost constraints for the DAQ.

cs.DC

Exploding AI Power Use: an Opportunity to Rethink Grid Planning and Management

The unprecedented rapid growth of computing demand for AI is projected to increase global annual datacenter (DC) growth from 7.2% to 11.3%. We project the 5-year AI DC demand for several power grids and assess whether they will allow desired AI growth (resource adequacy). If not, several "desperate measures" -- grid policies that enable more load growth and maintain grid reliability by sacrificing new DC reliability are considered. We find that two DC hotspots -- EirGrid (Ireland) and Dominion (US) -- will have difficulty accommodating new DCs needed by the AI growth. In EirGrid, relaxing new DC reliability guarantees increases the power available to 1.6x--4.1x while maintaining 99.6% actual power availability for the new DCs, sufficient for the 5-year AI demand. In Dominion, relaxing reliability guarantees increases available DC capacity similarly (1.5x--4.6x) but not enough for the 5-year AI demand. New DCs only receive 89% power availability. Study of other US power grids -- SPP, CAISO, ERCOT -- shows that sufficient capacity exists for the projected AI load growth. Our results suggest the need to rethink adequacy assessment and also grid planning and management. New research opportunities include coordinated planning, reliability models that incorporate load flexibility, and adaptive load abstractions.

cs.DC

Adapting Datacenter Capacity for Greener Datacenters and Grid

Cloud providers are adapting datacenter (DC) capacity to reduce carbon emissions. With hyperscale datacenters exceeding 100 MW individually, and in some grids exceeding 15% of power load, DC adaptation is large enough to harm power grid dynamics, increasing carbon emissions, power prices, or reduce grid reliability. To avoid harm, we explore coordination of DC capacity change varying scope in space and time. In space, coordination scope spans a single datacenter, a group of datacenters, and datacenters with the grid. In time, scope ranges from online to day-ahead. We also consider what DC and grid information is used (e.g. real-time and day-ahead average carbon, power price, and compute backlog). For example, in our proposed PlanShare scheme, each datacenter uses day-ahead information to create a capacity plan and shares it, allowing global grid optimization (over all loads, over entire day). We evaluate DC carbon emissions reduction. Results show that local coordination scope fails to reduce carbon emissions significantly (3.2%--5.4% reduction). Expanding coordination scope to a set of datacenters improves slightly (4.9%--7.3%). PlanShare, with grid-wide coordination and full-day capacity planning, performs the best. PlanShare reduces DC emissions by 11.6%--12.6%, 1.56x--1.26x better than the best local, online approach's results. PlanShare also achieves lower cost. We expect these advantages to increase as renewable generation in power grids increases. Further, a known full-day DC capacity plan provides a stable target for DC resource management.

cs.DC

Preparing for the Future -- Rethinking Proxy Apps

A considerable amount of research and engineering went into designing proxy applications, which represent common high-performance computing workloads, to co-design and evaluate the current generation of supercomputers, e.g., RIKEN's Supercomputer Fugaku, ANL's Aurora, or ORNL's Frontier. This process was necessary to standardize the procurement while avoiding duplicated effort at each HPC center to develop their own benchmarks. Unfortunately, proxy applications force HPC centers and providers (vendors) into a an undesirable state of rigidity, in contrast to the fast-moving trends of current technology and future heterogeneity. To accommodate an extremely-heterogeneous future, we have to reconsider how to co-design supercomputers during the next decade, and avoid repeating the past mistakes. This position paper outlines the current state-of-the-art in system co-design, challenges encountered over the past years, and a proposed plan to move forward.

cs.DC

Flexibility from Networks of Data Centers: A Market Clearing Formulation with Virtual Links

Data centers owned and operated by large companies have a high power consumption and this is expected to increase in the future. However, the ability to shift computing loads geographically and in time can provide flexibility to the power grid. We introduce the concept of virtual links to capture space-time load flexibility provided by geographically-distributed data centers in market clearing procedures. We show that the virtual link abstraction fits well into existing market clearing frameworks and can help analyze and establish market design properties. This is demonstrated using illustrative case studies.

eess.SY

Characterizing Curtailed and Uneconomic Renewable Power in the Mid-continent Independent System Operator

As power grids incorporate increased renewable generation such as wind and solar, their variability creates growing challenges for grid stability and efficiency. We study two facets: power the grid is unable to accept (curtailment), and power that is assigned zero economic value by the grid (negative or zero price). Collectively we term these "stranded power" or SP. We study stranded power in the Midcontinent Independent System Operator (MISO), characterizing quantity and temporal structure. First, stranded power is available in the MISO grid 99% of the time, and often in intervals >100 hours, with characteristic seasonal and time-of-day patterns. Average stranded power often exceeds 1 GW, with duty factors as high as 30%. About 30% of all wind generation in MISO is stranded. Examination of the top 10 individual sites shows stranded power can be as high as 70% duty factor and 250MW. Trends over the past 3.5 years suggest stranded power is a persistent phenomenon. The study characterizes opportunities to exploit stranded power. We consider using energy storage to increase the utility of stranded power. For a range of power levels and uniformly-distributed storage, adding 5 hours of storage doubles duty factor to 30% at 4MW, but another 95 hours is required for the next 15% increase. At 4MW with 50 hours of storage, only 3 of 200 sites reach 100% duty factor, and with 100 hours required for the next 10 sites. Higher power levels require 100's of hours. Storage at the top 10 sites is more productive, 5 hours increases duty factor to 70% at 4MW, but further storage has diminishing benefits. Studies of the amount of power served by storage show that distribution to the best sites provides 2 to 3.7-fold advantages over uniform distribution.

physics.soc-ph

Extreme Scaling of Supercomputing with Stranded Power: Costs and Capabilities

Power consumption (supply, heat, cost) and associated carbon emissions (environmental impact) are increasingly critical challenges in scaling supercomputing to Exascale and beyond. We proposes to exploit stranded power, renewable energy that has no value to the power grid, for scaling supercomputers, Zero-Carbon Cloud (ZCCloud), and showing that stranded power can be employed effectively to expand computing [1]. We build on those results with a new analysis of stranded power, characterizing temporal, geographic, and interval properties. We simulate production supercomputing workloads and model datacenter total-cost-of-ownership (TCO), assessing the costs and capabilities of stranded-power based supercomputing. Results show that the ZCCloud approach is cost-effective today in regions with high cost power. The ZCCloud approach reduces TCO by 21-45%, and improves cost-effectiveness up to 34%. We study many scenarios. With higher power price, cheaper computing hardware and higher system power density, benefits rise to 55%, 97% and 116% respectively. Finally, we study future extreme-scale systems, showing that beyond terascale, projected power requirements in excess of 100MW make ZCCloud up to 45% lower cost, for a fixed budget, increase peak PFLOPS achievable by 80%.

cs.DC

Data Centers as Dispatchable Loads to Harness Stranded Power

We analyze how both traditional data center integration and dispatchable load integration affect power grid efficiency. We use detailed network models, parallel optimization solvers, and thousands of renewable generation scenarios to perform our analysis. Our analysis reveals that significant spillage and stranded power will be observed in power grids as wind power levels are increased. A counter-intuitive finding is that collocating data centers with inflexible loads next to wind farms has limited impacts on renewable portfolio standard (RPS) goals because it provides limited system-level flexibility and can in fact increase stranded power and fossil-fueled generation. In contrast, optimally placing data centers that are dispatchable (with flexible loads) provides system-wide flexibility, reduces stranded power, and improves efficiency. In short, optimally placed dispatchable computing loads can enable better scaling to high RPS. We show that these dispatchable computing loads are powered to 60~80\% of their requested capacity, indicating that there are significant economic incentives provided by stranded power.

math.OC