Searcharxiv⌕ Search

arXiv · 2609.34048

Schedule Repair for DAG Workflows under Link Disruptions

Abstract

Schedules for directed acyclic graph (DAG) workflows in networked IoT systems are typically computed assuming a static or generally stable network. In contested and adversarial environments, this assumption is not valid. Links degrade and fail due to mobility, interference, and jamming. We study schedule repair: when a link disruption invalidates part of a schedule, how much of it should be rescheduled? We introduce a spectrum of repair policies that vary in repair scope, how much of the pending schedule each may move: wait out the disruption, reroute data around it, reschedule only the affected tasks locally, or reschedule all pending tasks globally. We evaluate each against an oracle and charge every repair a decision latency proportional to the extent to which it moves. Across 100 workload instances spanning synthetic task graphs, RIoTBench pipelines, and WfCommons scientific workflows, each run at five communication-to-computation ratios (CCRs) and disrupted by processes with deliberately different correlation structure, we find that no single scope wins: rerouting nearly erases isolated failures that cost waiting 30%, global repair comes within 4% of the oracle under jamming blackouts, waiting is favored under memoryless link flapping for larger and communication-heavy workloads (the scheduling analog of route-flap damping), self-healing mobility outages reward patience over reaction, and accounting for repair latency erodes large scopes first. We conclude that the scope of the repair should be adapted to the disruption process and the repair cost, rather than fixed by the scheduler.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mohammadali Khodabandehlou, Jared Coleman, Bhaskar Krishnamachari, Kevin Chan. 2026-09-28. Schedule Repair for DAG Workflows under Link Disruptions. https://arxiv.org/abs/2609.34048

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications

Large Language Models (LLMs) are increasingly deployed in complex multi-agent applications that rely on external function calls. This workload creates severe performance challenges for the KV Cache: spatial contention leads to the eviction of critical agents' caches and temporal underutilization leaves the cache of agents stalled on long-running function calls idling in GPU memory. We present TokenCake, a KV-Cache-centric serving framework that bridges this gap by co-optimizing scheduling and memory management through an agent-aware design. TokenCake's Temporal Scheduler employs an event-driven, opportunistic policy to proactively offload idle KV Caches during function calls and uses predictive uploading to hide data transfer latency. TokenCake's Spatial Scheduler uses dynamic memory partitioning, guided by a hybrid priority metric combining graph structure and runtime state, to reserve GPU memory for critical-path agents. Our evaluation on representative multi-agent benchmarks shows that TokenCake reduces end-to-end latency by up to 47.06% in memory-constrained settings and improves effective GPU utilization by 16.9 percentage points compared to vLLM. TokenCake is publicly available at https://github.com/pkulemonade/TokenCake.

cs.DC↗

SpSYRK: Half the Work in Distributed Sparse Matrix Multiplication

The symmetric rank-$k$ update (SYRK), $C = AA^\top$, computes the dot product of each pair of rows of $A$, producing the Gram matrix $C$. Its sparse variant underpins similarity search in machine learning, graph analytics, and genomics, including Jaccard similarity on datasets too large for a single node. Despite the symmetry in its inputs and outputs, existing distributed sparse matrix multiplication algorithms such as Sparse SUMMA treat sparse SYRK as generic multiplication, computing the full output and materializing the explicit transpose even when the calling application uses only one triangle. Prior distributed $AA^\top$ computations in similarity search and genome assembly inherit this overhead from the underlying SpGEMM. This paper presents SpSYRK and CommSpSYRK, two distributed sparse SYRK algorithms that exploit symmetry. The first, SpSYRK, partitions the off-diagonal blocks of the output between the upper and lower triangular regions of the process grid and computes only the lower-triangular part of each diagonal block, halving per-process computation compared with state-of-the-art distributed SpGEMM. The second, CommSpSYRK, further reorders communication to avoid forming $A^\top$, which reduces per-process communication volume. On 32 nodes of the Perlmutter supercomputer, SpSYRK achieves a 2$\times$ speedup over an optimized Sparse SUMMA on matrices where local multiplication dominates the runtime; the advantage narrows on communication-bound inputs, a dependence that the cost model predicts from the arithmetic intensity. CommSpSYRK fixes this and consistently achieves superior scaling at high process counts. The approach is a drop-in replacement for any application computing $C = AA^\top$ via a distributed SpGEMM routine, and its triangular output can be consumed directly by subsequent operations, reducing both computation and memory footprint.

cs.DC↗

Brain API: An Intent-Aware Control Plane for Policy-Governed Agentic Systems

Contemporary cloud and distributed systems expose control through resource-centric abstractions: services, deployments, network flows, execution graphs. Agentic and tool-augmented systems have meanwhile shifted application logic toward intent-driven, adaptive execution. Existing control planes, workflow engines and service meshes lack abstractions for intent-level decision governance: they cannot represent high-level goals as first-class control objects, cannot enforce policy over the mapping from intent to execution plan, and cannot produce auditable records of why one execution path was chosen over its alternatives. Control logic is therefore embedded in application code, leaving systems brittle, opaque and hard to govern. We propose Brain API, an intent-aware control plane for policy-governed agentic systems. Its central contribution is the decision artifact: a durable, versioned, auditable record of how an intent became an executable plan, capturing which policies applied, which capabilities were evaluated, which alternatives were rejected, and why. A motivating use case is agentic datasets: datasets participating as policy-governed capabilities under residency, compliance and cost constraints. We evaluate a prototype of the decision layer against two external policy corpora we did not author. On the OPA Gatekeeper constraint library it agrees with the library's own published verdicts on 42 of 42 encodable cases, 19 admit and 23 deny. On Cedar example policies, labeled by differential testing against its reference implementation, a deliberately dissimilar domain exposed three defects in our model, including a default-allow assumption that would have inverted every authorization policy. The evaluation covers policy filtering and selection; candidate generation, context signals, ranking and plan synthesis are not measured, nor is decision latency under load.

cs.DC↗