Searcharxiv⌕ Search

arXiv · 2609.39978

LatencyLab: A DPDK-Based P4 Pipeline Latency Measurement Framework for FPGA SmartNICs

Abstract

P4-programmable FPGA SmartNICs place packet processing directly on the wire, but open FPGA P4 toolflows do not expose timestamping at the pipeline boundary, so the latency a P4 program adds on the target FPGA is rarely measured. This paper presents LatencyLab, a DPDK-based measurement framework for FPGA P4 pipeline latency that needs neither PHC/PTP support on the datapath nor clock synchronization. The FPGA's two ports share a network segment, so the switch multicasts a copy of each probe packet to both: one copy passes through the VitisNetP4 pipeline, the other through a matched bypass path. A kernel-bypass DPDK receiver busy-polls both ports and timestamps every packet with the CPU timestamp counter (TSC) as it is retrieved from the NIC's receive circular buffer. The arrival-time difference of the two copies isolates the pipeline latency after calibration against a null bitstream carrying the same traffic; transmit time cancel in the subtraction. We evaluate four VitisNetP4 programs on an AMD Alveo U280, probing each with a 20,000-packet trace measured ten times per session over five independent sessions, all TSC-timestamped and reflected for hardware timestamping. The measured latency distributions are tight and reproducible: 99% of packets fall within 20 ns of the median, session medians repeating within 1 to 2 ns (FiveTuple 107/137 ns, Forward 149 ns, RemoveHeader 177 ns, Checksum 364 ns at 250 MHz). Two independent checks agree with the framework: a kernel-free reflector returns every probe pair to a ConnectX-5 NIC whose adapter clock reproduces the measured distributions within a few nanoseconds, quantile by quantile, and every measured packet falls 18 to 22 clock cycles below the vendor's worst-case latency bound.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pavani Kuppili, Zhaoyang Han, Yicheng Qian, Suranga Handagala, Michael Zink, Miriam Leeser, Robert Ricci. 2026-09-30. LatencyLab: A DPDK-Based P4 Pipeline Latency Measurement Framework for FPGA SmartNICs. https://arxiv.org/abs/2609.39978

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Deterministic Self-Stabilizing BFS Construction in Constant Space

In this paper, we resolve a long-standing question in self-stabilization by demonstrating that it is indeed possible to construct a spanning tree in a semi-uniform network using constant memory per node. We introduce a self-stabilizing synchronous algorithm that builds a breadth-first search (BFS) spanning tree with only $O(1)$ bits of memory per node, converging in $2^ε$ time units, where $ε$ denotes the eccentricity of the distinguish node. Crucially, our approach operates without any prior knowledge of global network parameters such as maximum degree, diameter, or total node count. In contrast to traditional self-stabilizing methods, such as pointer-to-neighbor communication or distance-to-root computation, that are unsuitable under strict memory constraints, our solution employs an innovative constant-space token dissemination mechanism. This mechanism effectively eliminates cycles and rectifies deviations in the BFS structure, ensuring both correctness and memory efficiency. The proposed algorithm not only meets the stringent requirements of memory-constrained distributed systems but also opens new avenues for research in self-stabilizing protocols under severe resource limitations.

cs.DC↗

Space Filling Curves is All You Need: Communication-Avoiding Matrix Multiplication Made Simple

General Matrix Multiplication (GEMM) is the cornerstone of HPC workloads and Deep Learning. State-of-the-art (SOTA) vendor libraries tune tensor layouts, parallelization schemes and cache blocking to minimize data movement across the memory hierarchy and maximize throughput. However, optimal settings for these parameters depend on the target platform and matrix shapes, making exhaustive tuning infeasible. In this work, we address this cumbersome scheduling search and tuning using space-filling curves (SFC). We partition the matrix multiplication using advancements in SFC, and obtain platform-oblivious and shape-oblivious matrix multiplication schemes with a high degree of data locality. We extend the SFC-based work partitioning to implement Communication-Avoiding (CA) algorithms with replication techniques in a seamless fashion. The resulting SFC-CA GEMM achieves provable asymptotic communication optimality for both square and rectangular matrix regimes. Across four x86 and Arm platforms, SFC-CA GEMM outperforms vendor libraries by up to 5.5$\times$ per shape and 1.8$\times$ in weighted harmonic mean (WHM) throughput. Last, we show the impact of our work on two real-world applications by leveraging our SFC-CA GEMM as a compute backend: i) prefill of LLM inference with speedups up to 1.85$\times$ over SOTA inference runtimes, and ii) distributed-memory matrix multiplication with speedups up to 2.3$\times$ over the SOTA distributed-memory GEMM framework with vendor-optimized compute backend.

cs.DC↗

MoEless: Efficient MoE LLM Serving with Serverless Experts

Large Language Models (LLMs) increasingly adopt Mixture-of-Experts (MoE) architectures to scale efficiently under stringent resource constraints. However, MoE's sparse activation causes severe expert load imbalance, where a few experts become stragglers while others remain underutilized, leading to inflated inference latency and cost. Existing solutions assume static, serverful model deployments, limiting expert elasticity and often incurring costly expert swapping or degraded output quality. We present MoEless, an efficient serverless MoE serving framework that mitigates expert load imbalance via elastic expert execution. MoEless leverages lightweight, layer-aware predictors to estimate incoming expert load distributions and proactively identify stragglers. We design optimized scaling and placement strategies to improve function locality, GPU utilization, and cross-expert load balance. MoEless is prototyped on top of Megatron-LM and deployed on an eight-GPU testbed. Experiments with open-source MoE models and real-world workloads show that MoEless reduces inference latency by 43% and inference cost by 84% compared to state-of-the-art solutions.

cs.DC↗