SearcharxivSearch

arXiv subjects

Kishwar Ahmed

Publications and source records attributed to Kishwar Ahmed.

5 recordsLinked to original sources

XOR Bidding and Knapsack Formulations for HPC Network Resource Allocation

Modern High Performance Computing (HPC) centers face growing challenges in ingesting large and diverse data streams. These issues often create bottlenecks that limit bandwidth utilization and delay scientific progress. Traditional static allocation and simple queuing methods are often insufficient. This paper presents a dynamic, value-based approach to bandwidth allocation. We formalize the problem by incorporating both network and processing constraints. To address it, we introduce two auction-based mechanisms: the Greedy Value Density Auction, which is computationally efficient, and the Vickrey--Clarke--Groves (VCG) Knapsack Auction, which provides strong theoretical guarantees. Both mechanisms rely on user bids that specify data requirements and scientific value. The objective is to maximize the total value of successful transfers, commonly referred to as social welfare. Simulation results demonstrate that the proposed mechanisms significantly outperform First Come First Served (FCFS) baselines. Under high-load conditions, they reduce average and tail completion delays by more than 80%. Predictability also improves, with the coefficient of variation of delay decreasing by 75--85%. Network stability increases as well, with load volatility, measured by the peak-to-average ratio, decreasing by 60--70%. These results indicate that value-driven, adaptive bandwidth allocation can reduce congestion, improve resource utilization, and provide fairer access based on scientific importance.

cs.NI

Power-Aware Scheduling for Multi-Center HPC Electricity Cost Optimization

This paper introduces TARDIS (Temporal Allocation for Resource Distribution using Intelligent Scheduling), a novel power-aware job scheduler for High-Performance Computing (HPC) systems that minimizes electricity costs through both temporal and spatial optimization. Our approach addresses the growing concerns of energy consumption in HPC centers, where electricity expenses constitute a substantial portion of operational costs and have a significant financial impact. TARDIS employs a Graph Neural Network (GNN) to accurately predict individual job power consumption, then uses these predictions to strategically schedule jobs across multiple HPC facilities based on time-varying electricity prices. The system integrates both temporal scheduling, shifting power-intensive workloads to off-peak hours, and spatial scheduling, distributing jobs across geographically dispersed centers with different pricing schemes. We evaluate TARDIS using trace-based simulations from real HPC workloads, demonstrating cost reductions of up to 18% in temporal optimization scenarios and 10 to 20% in multi-site environments compared to state-of-the-art scheduling approaches, while maintaining comparable system performance and job throughput. Our comprehensive evaluation shows that TARDIS effectively addresses limitations in existing power-aware scheduling approaches by combining accurate power prediction with holistic spatial-temporal optimization, providing a scalable solution for sustainable and cost-efficient HPC operations.

cs.DC

Scalable HPC Job Scheduling and Resource Management in SST

Efficient job scheduling and resource management contribute towards system throughput and efficiency maximization in high-performance computing (HPC) systems. In this paper, we introduce a scalable job scheduling and resource management component within the structural simulation toolkit (SST), a cycle-accurate and parallel discrete-event simulator. Our proposed simulator includes state-of-the-art job scheduling algorithms and resource management techniques. Additionally, it introduces workflow management components that support the simulation of task dependencies and resource allocations, crucial for workflows typical in scientific computing and data-intensive applications. We present the validation and scalability results of our job scheduling simulator. Simulation shows that our simulator achieves good accuracy in various metrics (e.g., job wait times, number of nodes usage) and also achieves good parallel performance.

cs.DC

HPC Application Parameter Autotuning on Edge Devices: A Bandit Learning Approach

The growing necessity for enhanced processing capabilities in edge devices with limited resources has led us to develop effective methods for improving high-performance computing (HPC) applications. In this paper, we introduce LASP (Lightweight Autotuning of Scientific Application Parameters), a novel strategy designed to address the parameter search space challenge in edge devices. Our strategy employs a multi-armed bandit (MAB) technique focused on online exploration and exploitation. Notably, LASP takes a dynamic approach, adapting seamlessly to changing environments. We tested LASP with four HPC applications: Lulesh, Kripke, Clomp, and Hypre. Its lightweight nature makes it particularly well-suited for resource-constrained edge devices. By employing the MAB framework to efficiently navigate the search space, we achieved significant performance improvements while adhering to the stringent computational limits of edge devices. Our experimental results demonstrate the effectiveness of LASP in optimizing parameter search on edge devices.

cs.PF

Practical Efficient Microservice Autoscaling with QoS Assurance

Cloud applications are increasingly moving away from monolithic services to agile microservices-based deployments. However, efficient resource management for microservices poses a significant hurdle due to the sheer number of loosely coupled and interacting components. The interdependencies between various microservices make existing cloud resource autoscaling techniques ineffective. Meanwhile, machine learning (ML) based approaches that try to capture the complex relationships in microservices require extensive training data and cause intentional SLO violations. Moreover, these ML-heavy approaches are slow in adapting to dynamically changing microservice operating environments. In this paper, we propose PEMA (Practical Efficient Microservice Autoscaling), a lightweight microservice resource manager that finds efficient resource allocation through opportunistic resource reduction. PEMA's lightweight design enables novel workload-aware and adaptive resource management. Using three prototype microservice implementations, we show that PEMA can find efficient resource allocation and save up to 33% resource compared to the commercial rule-based resource allocations.

cs.DC