Searcharxiv⌕ Search

arXiv · 2610.02498

Real-time Optimization of Simulation and Data Processing Pipelines for Experiments on Exascale Computing Platforms

Abstract

Current and next-generation large-scale experiments, especially in high-energy physics and cosmology, involve increasing volumes of data that require exascale computing power to analyze and model. Exascale machines present challenges for workflow scaling and efficient resource usage. We propose a multi-stage task scheduling approach for optimizing resource usage for these workflows. We demonstrate task scheduling for an example Monte Carlo simulation pipeline from the Short-Baseline Near Detector (SBND), and present theoretical and numerical tools for modeling task scheduling in these workflows that can be used to identify optimal scheduling strategies and choose resources for a given problem size.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Thomas Wester, Harikrishna Tummalapalli, Christine M. Simpson. 2026-10-01. Real-time Optimization of Simulation and Data Processing Pipelines for Experiments on Exascale Computing Platforms. https://arxiv.org/abs/2610.02498

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Grassroots Bonds: Financing by the People, for the People

Grassroots currencies turn mutual trust into liquidity: a grassroots coin is a unit of its issuer's debt, backed by the issuer's goods and services, which the issuer must redeem, 1-for-1, against any coin they hold; liquidity arises from mutual credit lines, formed by the voluntary exchange of coins among persons who trust each other. As coins are redeemable 1-for-1, the exchange must be 1-for-1 as well, lest prompt redemption after it leave one party with undue profit. Thus, grassroots coins are incongruent with interest-bearing credit. Here, we extend grassroots currencies to include also grassroots bonds, units of their issuer's debt due at a later date. Upon maturity, the bearer of the bond may redeem it against a coin of the issuer, so liquid coins can be lent against interest-bearing bonds. We specify bonds by adding Date and Escrow clauses to the social contract of grassroots currencies: the contract enforces the redemption of a bond when it is mature according to the date stated by the issuer, who undertakes to keep it current. We show that the voluntary swap of coins and bonds can realise the basic financial instruments: loans, sale of debt, and forward contracts, and, with an escrow agent, loans with payment schedules, options, collateral, guarantees, insurance, credit default swaps, letters of credit, and credit lines. We extend the liquidity ratios of corporate finance to bonds, with maturity as asset class, and prove that a community clears its debts by redemptions exactly when no member owes more than they hold, and that it does so without coordination once dates advance and persons act on their rights. Grassroots currencies that include coins and bonds are implemented in GLP, a concurrent logic programming language running on Dart, as a program derived from the contract and demonstrated by a village market of six agents and an escrow agent.

cs.DC↗

Characterization-Guided GPU Fault Resilience in NVIDIA MPS

NVIDIA Multi-Process Service (MPS) enables fine-grained GPU sharing by allowing multiple processes to execute concurrently on the same GPU, making it an important mechanism for improving GPU utilization. However, MPS has weak fault resilience: a fault in one process can terminate all co-running processes, limiting its adoption in resilience-critical settings such as multi-tenant GPU clusters. In this work, we design fault-resilient MPS to solve this problem. Our design is guided by insights from a systematic characterization of GPU faults and a deep analysis of their end-to-end processing pipeline. Based on these insights, we design two complementary mechanisms. First, we design a fault isolation mechanism for the dominant memory-related faults that can be fully isolated while preserving process-level fail-stop semantics by software intervention in the open GPU driver kernel module. For other faults whose process is within proprietary software, we design a fast-recovery substrate that combines virtual-memory-based GPU-resident state sharing with pre-initialized standbys. Our evaluation across GPUs and workloads demonstrates effective fault isolation and fast recovery with minimal overhead: in an end-to-end case study, isolation incurs no visible outage, while recovery restores pre-fault throughput in 355\,ms.

cs.DC↗

ARGOS: Reinforcement Learning-Driven Multidimensional Elasticity for Service Orchestration in the Computing Continuum

Data-intensive services in the Computing Continuum must balance analytics quality, resource usage, and cost across heterogeneous nodes with limited and uneven capacity. This balance becomes especially difficult when resource scaling reaches capacity limits, because changes in demand and cluster pressure must then be absorbed without violating client-defined quality ranges. Existing orchestrators mainly adapt resources, placements, or replicas, while analytics requirements such as coverage, sample, and freshness remain fixed. This article presents ARGOS, the Adaptive Reinforcement Learning-Driven Governance for Orchestrated Services, an end-to-end controller that formulates multidimensional elasticity as a per-request Markov decision process over analytics quality and cluster pressure, supported by capacity-aware admission. ARGOS is evaluated under controlled workloads and time-varying multi-tenant arrivals on a heterogeneous cluster. Across the controlled scenarios, the deep reinforcement learning policies consistently outperform the non-learning baselines and approach the independently tuned best-fixed reference. A separate live evaluation reports improvements over the static midpoint under realistic and saturated arrivals, with no recorded CPU or memory violations but remaining coverage violations. These results support deep reinforcement learning as an adaptive mechanism for multidimensional elasticity when resource scaling alone is insufficient.

cs.DC↗