Searcharxiv⌕ Search

arXiv · 2609.38235

Characterizing the eBPF-Based Data Plane for Multi-ClusterKubernetes: A Systematic Evaluation of Cilium Cluster Meshand KVStoreMesh

Abstract

Kubernetes deployments increasingly span multiple clusters for reasons of scale, fault isolation, regulatory boundaries, and heterogeneous hardware placement, including GPU-dense clusters dedicated to AI and HPC workloads. Cilium Cluster Mesh, built on extended Berkeley Packet Filter (eBPF) technology, has emerged as a widely deployed mechanism for connecting such clusters into a single logical network without a dedicated multi-cluster gateway. Despite its adoption, the academic literature contains no systematic, reproducible evaluation of the eBPF-based multi-cluster data plane itself: existing work either benchmarks eBPF against iptables within a single cluster, or applies eBPF to cross-cluster monitoring rather than forwarding and identity propagation. The most detailed scale data available - a report of a clustermesh-apiserver failing under load at approximately 45,000 nodes across 256 clusters - comes from an industry engineering blog rather than a peer-reviewed source. This paper proposes a methodology to close that gap. We define five research questions covering latency, throughput, control-plane limits, churn, and GPU/RDMA sensitivity, and describe a testbed and eBPF-level instrumentation design. Lacking a live multi-cluster testbed for this submission, we report results from a calibrated analytical simulation of control-plane architectures fit to the production data point. These results are presented as falsifiable predictions intended for validation once the testbed is complete. This paper provides the motivation, experimental design, and preliminary simulation results to serve as a template for a full empirical study.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Simhadri Podala Narasimha. 2026-09-28. Characterizing the eBPF-Based Data Plane for Multi-ClusterKubernetes: A Systematic Evaluation of Cilium Cluster Meshand KVStoreMesh. https://doi.org/10.48550/arxiv.2609.38235

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

End-to-end Modeling and Optimization of Timing and Energy in Wireless IoT Systems

With the advent of edge computing, data generated by end devices can be pre-processed before transmission, possibly saving transmission time and energy. On the other hand, data processing itself incurs latency and energy consumption, depending on the complexity of the computing operations and the speed of the processor. The energy-latency-reliability profile resulting from the concatenation of pre-processing operations (specifically, data compression) and data transmission is particularly relevant in wireless communication services, whose requirements may change dramatically with the application domain. In this paper, we study this multi-dimensional optimization problem, introducing a simple model to investigate the tradeoff among end-to-end latency, reliability, and energy consumption when considering compression and communication operations in a constrained wireless device. We then study the Pareto fronts of the energy-latency trade-off, considering data compression ratio and device processing speed as key design variables. Our results show that the energy costs grows exponentially with the reduction of the end-to-end latency, so that considerable energy saving can be obtained by slightly relaxing the latency requirements of applications. These findings challenge conventional rigid communication latency targets, advocating instead for application-specific end-to-end latency budgets that account for computational and transmission overhead.

cs.NI↗

Optimal Oblivious Load-Balancing for Sparse Traffic in Large-Scale Satellite Networks

Oblivious load-balancing in networks involves routing traffic from sources to destinations using predetermined routes independent of the traffic, so that the maximum load on any link in the network is minimized. We investigate oblivious load-balancing schemes for a $N\times N$ torus network under sparse traffic where there are at most $k$ active source-destination pairs. We are motivated by the problem of load-balancing in large-scale LEO satellite networks, which can be modelled as a torus, where the traffic is known to be sparse and localized to certain hotspot areas. We formulate the problem as a linear program and show that no oblivious routing scheme can achieve a worst-case load lower than approximately $\frac{\sqrt{2k}}{4}$ when $1<k \leq N^2/2$ and $\frac{N}{4}$ when $N^2/2\leq k\leq N^2$. Moreover, we demonstrate that the celebrated Valiant Load Balancing scheme is suboptimal under sparse traffic and construct an optimal oblivious load-balancing scheme that achieves the lower bound. Further, we discover a $\sqrt{2}$ multiplicative gap between the worst-case load of a non-oblivious routing and the worst-case load of any oblivious routing. The results can also be extended to general $N\times M$ tori with unequal link capacities along the vertical and horizontal directions.

cs.NI↗

SAFE: Spatially-Aware Feedback Enhancement for Fault-Tolerant Trust Management in Event-Based VANETs

In event-based trust management for vehicular ad hoc networks (VANETs), vehicles that witness a road event broadcast its state, and vehicles that later witness the same event score the earlier broadcasters in feedback reports sent to a central decision unit (CDU). When the event state changes, an honest vehicle that reported the previous state receives negative feedback, as the evaluator compares an out-of-date message with the new state. This stale-witness problem is hypothesised to be a major source of unfair penalisation of honest vehicles. To address it, SAFE (Spatially-Aware Feedback Enhancement) is proposed. In SAFE, vehicles continue to record event messages after their decision and throughout the witness area, and send an updated feedback report when they leave it. SAFE was compared with the trust cascading-based emergency message dissemination model (TCEMD) in attack-free highway scenarios simulated with OMNeT++, Veins and Simulation of Urban MObility (SUMO). In the single-event scenario, the negative-feedback rate in the two update rounds after the state change decreased from 53.8% to 24.7% and from 55.6% to 8.3%, and the number of distinct honest vehicles blacklisted decreased from 20 to 6. In the multi-event scenario, the negative-feedback rate remained at or below 1.3% in SAFE, compared with up to 77.6% in TCEMD, and the share of trust evaluations ending in an untrusted label decreased from 22.7% to at most 1.7%. These gains required 2.3 to 5.0 times more feedback entries. Experiments with two decision distances showed that a shorter gap between the decision and witness distances reduced stale feedback in both schemes, supporting the stale-witness hypothesis

cs.NI↗