arXiv · 2607.25563
Few-Shot Open-Vocabulary Remote Sensing Segmentation via Textual Inversion
Abstract
Open-vocabulary segmentation labels arbitrary categories from a text query without per-class training, yet on remote sensing imagery it underperforms on categories it handles reliably elsewhere. We find that much of this gap traces to the text query rather than to the segmentation model. Because these models are not specialized for overhead imagery, the class name that serves as the query is often a weak address into the vision-language embedding space. We show that a better name repairs part of the gap, while the remaining failures call for an address that the tested natural-language rephrasings do not provide. We recover that address from a few examples through textual inversion on a frozen model, keeping inference text only. On a representative benchmark this raises the mean intersection over union on the affected categories from 3.9 to 39.4, and across eight remote sensing datasets it improves over few-shot methods that instead inject visual prompts at inference.
Explore related subjects
Keep this discovery
Junhyuk Heo, Junghwan Park. 2026-07-28. Few-Shot Open-Vocabulary Remote Sensing Segmentation via Textual Inversion. https://arxiv.org/abs/2607.25563
Cite the original work for its findings. Save a collection to share your selection of sources.