arXiv · 2603.27241
SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track
Abstract
Referring video object segmentation (RVOS) commonly grounds targets in videos based on static textual cues. MeViS benchmark extends this by incorporating motion-centric expressions (referring & reasoning motion expressions) and introducing no-target queries. Extending SaSaSa2VA, where increased input frames and [SEG] tokens already strengthen the Sa2VA backbone, we adopt a simple yet effective target existence-aware verification mechanism, leading to Still Awesome SaSaSa2VA (SaSaSaSa2VA). Despite its simplicity, the method achieves a final score of 89.19 in the 5th PVUW Challenge (MeViS-Text Track), securing 2nd place. Both quantitative results and ablations suggest that this existence-aware verification strategy is sufficient to unlock strong performance on motion-centric referring tasks.
Explore related subjects
Keep this discovery
Dengxian Gong, Quanzhu Niu, Shihao Chen, Yuanzheng Wu, Yikang Zhou, Tao Zhang, Haobo Yuan, Lu Qi, Shunping Ji. 2026-03-28. SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track. https://arxiv.org/abs/2603.27241
Cite the original work for its findings. Save a collection to share your selection of sources.