SearcharxivSearch

arXiv subjects

Yansong Pei

Publications and source records attributed to Yansong Pei.

5 recordsLinked to original sources

Real-Time Edge-based Detection of Correlated AI Data-Center Load Episodes

Artificial-intelligence data centers running bulk-synchronous training can impose sub-second power swings. When several facilities synchronize their training cycles, these load variations become spatially correlated and amplify the aggregate disturbance on the grid. A grid operator without access to data-center telemetry must infer this correlation from electrical measurements alone. However, the required observation time and the feasibility of detection on substation-deployable hardware remain uncharacterized. This paper develops a correlation-based detection method to classify the multi-facility operating regime from cross-facility power measurements. Analytical derivations and experimental validation show that the resulting detection confidence increases with the observation-window length at a rate governed by the load correlation time. The method is demonstrated in a real-time hardware-in-the-loop testbed, where load setpoints generated from a validated semi-Markov data-center load model are applied to an electromagnetic-transient grid simulation on a Real-Time Digital Simulator. A compact classifier built on pairwise power correlations runs on an edge device in this loop and determines whether the data-center load variations are independent or spatially correlated. The cross-facility correlation separates the independent and correlated cases across independent realizations. The held-out detection accuracy improves with the observation window, consistent with the predicted relation. A raw-waveform network fails to generalize, supporting pairwise correlation as the discriminative signal. The detector executes in real time on commodity edge hardware. A closed-loop demonstration against the running simulator tracks a regime change within one observation window.

eess.SY

Detection of Synchronized AI Data Center Load Episodes Using SCADA Telemetry

AI data centers running distributed training workloads impose episodic, spatially correlated active-power disturbances on the transmission grid. These synchronized episodes increase cross-substation load correlation and limit the diversification benefit that reserve margin planning assumes. Conventional energy management systems evaluate each substation independently and do not extract the cross-substation statistical structure that defines a synchronized episode. This paper develops a detection method that identifies synchronized AI data-center load episodes from standard active-power telemetry without new instrumentation, trained classifiers, or labeled data. The method computes a Synchronization Index, the dominant eigenvalue fraction of a sliding sample covariance matrix across substations. A cumulative-sum (CUSUM) sequential test converts the index into a delay-bounded episode alarm. The same eigen decomposition yields, at no additional cost, a dominant eigenvector that attributes a detected episode to the substations that drive it. Tests on a real-time digital simulator (RTDS) model of the IEEE 39-bus system with three AI data-center buses show that the method separates episode and normal windows with a wide margin over chance and attributes episode participation at substation granularity.

eess.SY

Storage as a Transmission Asset (SATA) for Large-Load Congestion Relief

Hyperscale data centers and other large concentrated loads can impose substantial new demand on existing transmission networks. If import corridors lack sufficient transfer capability, operators may need to curtail load, delay interconnection, or reinforce the network to maintain reliable service. An energy storage system (ESS) deployed as a storage-as-transmission asset (SATA) offers a non-wires alternative by providing operator-directed support to constrained import corridors. However, the operating-level reliability value of SATA dispatch remains insufficiently quantified. This paper evaluates operator-directed SATA using a day-ahead DC optimal power flow that co-optimizes generation, ESS dispatch, and load curtailment across Monte Carlo scenarios of demand and generator availability. Operating reliability is assessed using expected energy not served (EENS), loss-of-load hours (LOLH), and the conditional value at risk (CVaR) of daily unserved energy. Congestion-price and flow-sensitivity metrics are used to identify the limiting corridor and storage location. The interconnection is then screened to determine whether SATA is suitable, reinforcement is required, or storage would provide little transmission value. Results show that operator-directed SATA reduces average unserved energy, loss-of-load exposure, and tail risk compared with deploying the same ESS for pure arbitrage. These results demonstrate that the operating designation of storage is a primary driver of its transmission value.

eess.SY

A Pre-Dispatch Resonance Safety Criterion for AI Training Clusters

Hyperscale AI training clusters operate under the Bulk Synchronous Parallel protocol, which impose a periodic power swing on the transmission grid. Every GPU in the job transitions between compute and idle in lockstep, so the aggregate power traces a square wave at the training iteration period. Production iteration periods of one to ten seconds place the forcing frequency within the inter-area electromechanical mode band of large interconnections, where a training schedule can drive a mode at resonance. This paper derives a closed-form pre-dispatch safety criterion that bounds the maximum cluster size a grid can absorb at any proposed iteration period. The derivation inverts the steady-state forced two-area swing equations. The criterion defines a danger band of iteration periods, extends to the square-wave harmonics, and parameterizes the modal response from planning-study eigenanalysis and the forcing amplitude from GPU specifications. Applied to the IEEE 39-bus system at a production-representative duty cycle, the criterion shows that the maximum safe cluster at resonance is $66\,900$ GPUs under light damping. Rescheduling the same job less than one second away from resonance reduces the deviation $7.4\times$ with no hardware change. These results establish the training iteration period as a controllable grid-safety parameter and supply the analytic screening tool that reliability directives on current large loads lack.

eess.SY

Spatial Load Correlation in AI Data-Center-Dominated Power Systems

The proliferation of large-scale data centers introduces spatially correlated demand profiles that challenge the long-standing assumption of statistical independence of loads in power system analysis. This paper examines the emergence of such load correlations and evaluates their impact on data-center-dominated grids. Analytical derivations reveal that correlated load fluctuations amplify aggregate stochastic disturbances, reduce voltage stability margins through weakened reactive power stiffness, and degrade frequency stability margin by erosion of natural load diversity effects. Real-time digital simulation studies confirm that moderate spatial correlation in distributed data centers produces simultaneous frequency deviations and voltage fluctuations across multiple buses. The findings offer transmission system operators a physics-based perspective to interpret emerging oscillatory phenomena and establish stability planning criteria grounded in measurable load-correlation structures rather than traditional diversity assumptions.

eess.SY