Searcharxiv⌕ Search

arXiv subjects

Linlin Han

Publications and source records attributed to Linlin Han.

1 recordsLinked to original sources

From Noisy Telemetry to Actionable Warnings: GPU Failure Prediction in Industrial Clusters

GPU clusters are critical infrastructure for AI services, but accurate and actionable GPU failure prediction remains a problem in production settings. We study ticket-linked telemetry from a ByteDance GPU cluster and identify three obstacles: workload-confounded telemetry, heterogeneous fault precursors, and the gap between window-level predictions and actionable alerts. These findings motivate Falcon, a fault-specific warning framework combining missingness-aware temporal and peer-relative features, fault-specific learner selection, and an event policy based on thresholding, persistence, and cooldown. On the test set, Falcon achieves the highest F1 among four baselines and reaches 70.6% F1 on the best-performing fault type. Detected cases provide median lead times of 17.34-35.57 hours. We further report a production deployment, where Falcon is calibrated toward high-precision alerts to reflect false-positive costs. Together, these results show that fault-specific modeling improves early warning from noisy production GPU telemetry.

cs.SE↗