arXiv · 2604.23554
Adaptive Swin Transformer Partitioning over AI-RAN Networks
Abstract
This paper demonstrates the feasibility of transformer-based split inference for real-time video object detection over dynamic 5G AI-RAN networks. We extend throughput-aware adaptive splitting from CNNs to a Swin Transformer backbone and show that practical split execution is achievable for transformer-based vision models without retraining. To address the large intermediate activations inherent to transformers, we introduce an efficient, accuracy-preserving activation compression pipeline that substantially reduces uplink payload. The complete system -- including adaptive split selection, transformer inference, and compression -- is implemented and validated end-to-end on a real-time detection workload, with distributed UPF (dUPF) integration further reducing user-plane latency and improving runtime stability. Extensive measurements on an NVIDIA Aerial-based AI-RAN testbed jointly account for inference and 5G communication energy, quantifying the latency-energy-privacy trade-offs in realistic deployments.
Explore related subjects
Keep this discovery
Tam Thanh Nguyen, Yong Hao Pua, Tuan Van Ngo, Mao V. Ngo, Jihong Park, Binbin Chen, Tony Q. S. Quek. 2026-04-26. Adaptive Swin Transformer Partitioning over AI-RAN Networks. https://arxiv.org/abs/2604.23554
Cite the original work for its findings. Save a collection to share your selection of sources.