SearcharxivSearch

arXiv subjects

Vasilis Bountris

Publications and source records attributed to Vasilis Bountris.

9 recordsLinked to original sources

How Accurately Can the Energy Use of Spark Applications Be Estimated Based on Resource Utilisation?

Distributed batch data processing applications are widely executed on cloud-based resources where restricted user access to node-level hardware energy counters hinders transparent sustainability accounting. Energy and carbon attribution methodologies therefore depend on power models and available resource utilisation traces, yet the accuracy of these estimates has to be validated while direct counters are available. In this work, we use Apache Spark running on Kubernetes as a case-study dataflow runtime and cluster resource manager to compare model-based energy estimates to Intel RAPL package and DRAM energy on an AWS bare-metal cloud and an on-premises cluster, comparing different CPU usage signals and memory coefficients. We show that external monitoring improves signed package-energy error relative to Spark task traces, reducing underestimation from -29.58% to -24.41% on AWS and from -24.00% to -16.22% on-premises.

cs.DC

Ichnos+: Estimating the Carbon Footprint of Scientific Workflows Using Fitted Power Models

As data-intensive scientific workflows scale to facilitate the automation of analysis of increasing amounts of data, their resource-intensive and long-running execution incurs significant energy consumption and carbon emissions. Given the already significant and rising emissions from the ICT sector, it is crucial to quantify and understand the carbon footprint of scientific workflows. However, existing tooling is commonly not usable in shared, virtualized environments or resorts to power models that are based on only one or two generic data points. To address this gap, this paper presents Ichnos+, a novel system to quantify the environmental footprint of Nextflow scientific workflows. Ichnos+ enables post-hoc footprint estimation based on existing workflow traces, node-specific power models for the computational resources utilized, and carbon intensity data aligned with the execution time. We evaluate Ichnos+ against hardware-level energy measurements obtained using Intel RAPL, and the nf-core co2footprint plugin, which implements the Green Algorithms methodology. We find that Ichnos+ is capable of estimating workflow energy consumption with an estimation error of 10.8% across three compute clusters, significantly outperforming the nf-core plugin. We further show that Ichnos+ extends beyond operational carbon to estimate embodied emissions as well as water and land use. Finally, we demonstrate how Ichnos+ can be extended for another workflow system, Apache Airflow, maintaining a similarly high degree of estimation accuracy.

cs.DC

Augur: Pre-Execution Energy Prediction for Workflow Tasks in Heterogeneous Clusters

Scientific workflows are widely used to process large quantities of data, leading to significant energy consumption and carbon emissions. To reduce this environmental impact, energy and carbon-aware scheduling approaches could be employed. However, such methods require runtime and energy predictions, which are typically only available for workflows that have been executed previously. Meanwhile, scientists may execute new or modified workflows, use workflows with different input data, or run them on alternative infrastructure. To address this critical gap, we propose Augur, a novel method to predict the energy consumption of scientific workflow tasks prior to execution. By efficiently profiling both the available cluster infrastructure and the workflow at hand, Augur is capable of predicting the overall energy consumption of the workflow with a median prediction error of $16.3\pm15.3\%$ compared to Ichnos, an energy estimation method that uses fitted power models, and $18.2\pm14.7\%$ compared to Intel RAPL, as observed in our experimental evaluation on public and private cloud infrastructure. Relying on only minimal historical execution data, Augur outperforms two state-of-the-art methods in predicting both task runtime and total workflow energy, providing a robust foundation for energy-efficient and carbon-aware scientific data analysis.

cs.DC

A Systematic Evaluation of the Potential of Carbon-Aware Execution for Scientific Workflows

Scientific workflows are critical to scientific data analysis and often involve computationally intensive processing of large datasets on compute clusters. As such, their execution tends to be long-running and resource-intensive, resulting in significant energy consumption and carbon emissions. While carbon-aware computing methods have received considerable attention in general cloud contexts, their application to scientific data analysis workflows remains a critical research gap. Our study addresses this oversight by showing how the delay tolerance, interruptibility, and scalability of scientific workflows can be leveraged for a significantly more sustainable execution model. In this study, we first quantify the problem of carbon emissions associated with running scientific workflows, and then demonstrate the transformative potential for carbon-aware workflow execution. We estimate the carbon footprint of seven real-world Nextflow workflows executed on diverse dedicated cluster and public cloud resources using high-resolution average and marginal grid carbon intensity data from open and commercial data providers. Furthermore, we conduct a systematic evaluation of the impact of carbon-aware temporal shifting, and the dynamic pausing and resuming of the workflow. Moreover, we investigate the impact of resource scaling at both workflow and workflow task levels. Finally, we report substantial potential reductions in overall carbon emissions, with temporal shifting capable of decreasing emissions by over 80%, and resource scaling by 67%.

cs.DC

HyProv: Hybrid Provenance Management for Scientific Workflows

Provenance plays a crucial role in scientific workflow execution, for instance by providing data for failure analysis, real-time monitoring, or statistics on resource utilization for right-sizing allocations. The workflows themselves, however, become increasingly complex in terms of involved components. Furthermore, they are executed on distributed cluster infrastructures, which makes the real-time collection, integration, and analysis of provenance data challenging. Existing provenance systems struggle to balance scalability, real-time processing, online provenance analytics, and integration across different components and compute resources. Moreover, most provenance solutions are not workflow-aware; by focusing on arbitrary workloads, they miss opportunities for workflow systems where optimization and analysis can exploit the availability of a workflow specification that dictates, to some degree, task execution orders and provides abstractions for physical tasks at a logical level. In this paper, we present HyProv, a hybrid provenance management system that combines centralized and federated paradigms to offer scalable, online, and workflow-aware queries over workflow provenance traces. HyProv uses a centralized component for efficient management of the small and stable workflow-specification-specific provenance, and complements this with federated querying over different scalable monitoring and provenance databases for the large-scale execution logs. This enables low-latency access to current execution data. Furthermore, the design supports complex provenance queries, which we exemplify for the workflow system Airflow in combination with the resource manager Kubernetes. Our experiments indicate that HyProv scales to large workflows, answers provenance queries with sub-second latencies, and adds only modest CPU and memory overhead to the cluster.

cs.DC

Exploring the Potential of Carbon-Aware Execution for Scientific Workflows

Scientific workflows are widely used to automate scientific data analysis and often involve processing large quantities of data on compute clusters. As such, their execution tends to be long-running and resource intensive, leading to significant energy consumption and carbon emissions. Meanwhile, a wealth of carbon-aware computing methods have been proposed, yet little work has focused specifically on scientific workflows, even though they present a substantial opportunity for carbon-aware computing because they are inherently delay tolerant, efficiently interruptible, and highly scalable. In this study, we demonstrate the potential for carbon-aware workflow execution. For this, we estimate the carbon footprint of two real-world Nextflow workflows executed on cluster infrastructure. We use a linear power model for energy consumption estimates and real-world average and marginal CI data for two regions. We evaluate the impact of carbon-aware temporal shifting, pausing and resuming, and resource scaling. Our findings highlight significant potential for reducing emissions of workflows and workflow tasks.

cs.DC

Ponder: Online Prediction of Task Memory Requirements for Scientific Workflows

Scientific workflows are used to analyze large amounts of data. These workflows comprise numerous tasks, many of which are executed repeatedly, running the same custom program on different inputs. Users specify resource allocations for each task, which must be sufficient for all inputs to prevent task failures. As a result, task memory allocations tend to be overly conservative, wasting precious cluster resources, limiting overall parallelism, and increasing workflow makespan. In this paper, we first benchmark a state-of-the-art method on four real-life workflows from the nf-core workflow repository. This analysis reveals that certain assumptions underlying current prediction methods, which typically were evaluated only on simulated workflows, cannot generally be confirmed for real workflows and executions. We then present Ponder, a new online task-sizing strategy that considers and chooses between different methods to cater to different memory demand patterns. We implemented Ponder for Nextflow and made the code publicly available. In an experimental evaluation that also considers the impact of memory predictions on scheduling, Ponder improves Memory Allocation Quality on average by 71.0% and makespan by 21.8% in comparison to a state-of-the-art method. Moreover, Ponder produces 93.8% fewer task failures.

cs.DC

Low-level I/O Monitoring for Scientific Workflows

While detailed resource usage monitoring is possible on the low-level using proper tools, associating such usage with higher-level abstractions in the application layer that actually cause the resource usage in the first place presents a number of challenges. Suppose a large-scale scientific data analysis workflow is run using a distributed execution environment such as a compute cluster or cloud environment and we want to analyze the I/O behaviour of it to find and alleviate potential bottlenecks. Different tasks of the workflow can be assigned to arbitrary compute nodes and may even share the same compute nodes. Thus, locally observed resource usage is not directly associated with the individual workflow tasks. By acquiring resource usage profiles of the involved nodes, we seek to correlate the trace data to the workflow and its individual tasks. To accomplish that, we select the proper set of metadata associated with low-level traces that let us associate them with higher-level task information obtained from log files of the workflow execution as well as the job management using a task orchestrator such as Kubernetes with its container management. Ensuring a proper information chain allows the classification of observed I/O on a logical task level and may reveal the most costly or inefficient tasks of a scientific workflow that are most promising for optimization.

cs.DC

Large Language Models to the Rescue: Reducing the Complexity in Scientific Workflow Development Using ChatGPT

Scientific workflow systems are increasingly popular for expressing and executing complex data analysis pipelines over large datasets, as they offer reproducibility, dependability, and scalability of analyses by automatic parallelization on large compute clusters. However, implementing workflows is difficult due to the involvement of many black-box tools and the deep infrastructure stack necessary for their execution. Simultaneously, user-supporting tools are rare, and the number of available examples is much lower than in classical programming languages. To address these challenges, we investigate the efficiency of Large Language Models (LLMs), specifically ChatGPT, to support users when dealing with scientific workflows. We performed three user studies in two scientific domains to evaluate ChatGPT for comprehending, adapting, and extending workflows. Our results indicate that LLMs efficiently interpret workflows but achieve lower performance for exchanging components or purposeful workflow extensions. We characterize their limitations in these challenging scenarios and suggest future research directions.

cs.DC