SearcharxivSearch

arXiv subjects

Thomas Magnanti

Publications and source records attributed to Thomas Magnanti.

3 recordsLinked to original sources

Refined Thompson Learning for Adaptive Bandits: Power-Efficient Flexibility Scheduling Across Data Centers

The rapid growth of large-scale AI workloads in data centers has placed increasing pressure on power grids in recent years. Since power systems must continuously balance supply and demand, there is growing interests in leveraging data-center workload flexibility as a grid service. We propose a contextual restless multi-armed bandit (CRMAB) framework in which a grid operator requests load reductions without observing internal job-scheduling decisions. Under index-ability guarantee, each data center or physical machine is modeled as a Markov decision process (MDP) over a cyclic virtual-machine (VM) job queue, with unknown rewards and transition dynamics learned online using Thompson sampling and Whittle-index policies. To improve learning under sparse and noisy observations, the framework augments an adaptive Thompson--Whittle (TW) policy with domain-informed transition priors and gated prior mixing. In baseline experiments, the best adaptive refined variant achieves 91.4\% of the oracle reward after 100 rounds and 96.8\% after 1,000 rounds. Across a 16-setting stress test spanning different state-space sizes and levels of contextual noise, the best refined variant consistently outperforms the original TW policy with high confidence while remaining competitive with EXP4. A graph-based prior further incorporates data-center hardware constraints, including computing-resource limits. Overall, the results demonstrate the economic potential of data-center flexibility as a grid service and highlight the importance of high-quality, open-source AI workload traces for developing and evaluating such services.

cs.CE

Robust Restless Multi-Armed Bandit for Data Center Flexibility Services Through Virtual Machine Scheduling

Energy demands from data centers have surged and stressed the grid in recent years. Electric grids require balancing supply and demand every second, motivating demand response (reduction) from large loads, including data centers. This can be achieved by rescheduling jobs on a physical machine. Its real-time implementation is uncertain due to fluctuating resource utilization, and rescheduling incurs quality-of-service (QoS) losses that providers are unwilling to disclose. We propose a restless multi-armed bandit (RMAB) framework, in which the grid operator requests load reductions without access to detailed job-rescheduling procedures. Using open-source virtual machine (VM) datasets, we model job arrivals and rescheduling at each data center as a restless arm in a Markov decision process (MDP) and derive Whittle-index-based policies using the learned transition function via Thompson sampling. To overcome the weakness of an increasingly long learning process due to an enlarged state space, we use a mixed strategy that includes a global upper confidence bound (UCB) and encodes trust indices to enhance robustness and accelerate learning. Results show that the proposed mixed-strategy algorithm remains robust across varying state-space sizes and consistently outperforms the pure Thompson-Whittle (TW) algorithm, especially when contextual information is noisy. It also demonstrates superior performance compared to the state-of-the-art EXP4 framework. We provided open-source code to ensure reproducibility.

cs.CE

Decision-dependent Robust Charging Infrastructure Planning for Light-duty Truck Electrification at Industrial Sites

Many industrial sites and digital logistics platforms rely on diesel-powered light-duty trucks to transport workers and small-scale facilities, which results in a significant amount of greenhouse gas emissions (GHGs). To address this, we develop a robust model for planning charging infrastructure to electrify light-duty trucks at industrial sites. The model is formulated as a mixed-integer linear program (MILP) that optimizes the charging infrastructure selection (across multiple charger types and locations) and determines charging schedules for each truck based on the selected infrastructure. Given the strict stop times and schedules at industrial sites, we introduce a scheduling-with-abandonment problem in which trucks forgo charging if their waiting time exceeds a maximum threshold. We further incorporate the impacts of overnight charging and range anxiety on drivers' waiting and abandonment behaviors. To model stochastic, heterogeneous parking durations, we classified trucks using machine learning (ML) methods based on contextual and time-location features. We then constructed decision-dependent, feature-driven robust uncertainty sets in which parking-time variability varies flexibly with drivers' charging choices. These feature-driven sets are applied to two robust optimization formulations with decision-dependent uncertainty (RO-DDU), resulting in distinct outcomes and managerial implications. We conduct a case study at an open-pit mining site to plan charger installations across eight charging zones, serving approximately 200 trucks. By decomposing the problem into a short rolling horizon or using a heuristic approach for the full-year or representative-day dataset, the model achieves an optimality gap of less than 0.1\% under diverse uncertainty scenarios.

eess.SY