arXiv · 2410.12979
BlabberSeg: Real-Time Embedded Open-Vocabulary Aerial Segmentation
Abstract
Real-time aerial image segmentation plays an important role in the environmental perception of Uncrewed Aerial Vehicles (UAVs). We introduce BlabberSeg, an optimized Vision-Language Model built on CLIPSeg for on-board, real-time processing of aerial images by UAVs. BlabberSeg improves the efficiency of CLIPSeg by reusing prompt and model features, reducing computational overhead while achieving real-time open-vocabulary aerial segmentation. We validated BlabberSeg in a safe landing scenario using the Dynamic Open-Vocabulary Enhanced SafE-Landing with Intelligence (DOVESEI) framework, which uses visual servoing and open-vocabulary segmentation. BlabberSeg reduces computational costs significantly, with a speed increase of 927.41% (16.78 Hz) on a NVIDIA Jetson Orin AGX (64GB) compared with the original CLIPSeg (1.81Hz), achieving real-time aerial segmentation with negligible loss in accuracy (2.1% as the ratio of the correctly segmented area with respect to CLIPSeg). BlabberSeg's source code is open and available online.
Explore related subjects
Keep this discovery
Haechan Mark Bong, Ricardo de Azambuja, Giovanni Beltrame. 2024-10-16. BlabberSeg: Real-Time Embedded Open-Vocabulary Aerial Segmentation. https://arxiv.org/abs/2410.12979
Cite the original work for its findings. Save a collection to share your selection of sources.