arXiv · 2609.25022
NPLSD: Accelerating Line-Segment Detection on NPU Microcontrollers
Abstract
Line-segment detection is fundamental to robotics, autonomous navigation, and industrial inspection. While transformer-based detectors achieve the highest accuracy, their deployment on microcontrollers remains impractical due to resource constraints. The STM32N6, with its Neural-ART NPU, promises to enable deep vision at the extreme edge. However, existing detectors rely on attention, grid-sampling, and normalization, operators that are unsupported by the convolution-oriented NPU. This architectural mismatch is characterized operator by operator: attention, grid-sampling, and normalization lack accelerator primitives, and the decoder's self-attention alone materializes a 39 MB score tensor that exceeds on-chip memory. To address this limitation, NPLSD is introduced as a pair of NPU-compatible line-segment detectors built from one design methodology. NPLSD-H retains the convolutional HGNetv2 backbone of LINEA and replaces the transformer head with a fully-convolutional feature pyramid and an F-Clip dense head. NPLSD-M adapts the M-LSD-tiny trunk to the supported operator set. Warm-started from ImageNet and trained on ShanghaiTech Wireframe, the 2.63M-parameter NPLSD-H reaches sAP^10=37.9 (35.9 int8); the 0.62M-parameter NPLSD-M reaches 41.9 (41.1 int8). A controlled ablation isolates the trunk as the only variable, and initialization alone accounts for 4.6 points.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Parsa Hassani Shariat Panahi, Amir Hossein Jalilvand, M. Hassan Najafi. 2026-08-12. NPLSD: Accelerating Line-Segment Detection on NPU Microcontrollers. https://arxiv.org/abs/2609.25022
Cite the original work for its findings. Save a collection to share your selection of sources.