QST-SAM: Leveraging Cross-modal Instructions for Few-shot Referring Video Object Segmentation
Abstract
The advancement of Video Object Segmentation (VOS) isimpeded by its heavy reliance on expensive annotations. This has led toa landscape of compromised alternatives: unsupervised methods provideweak information, few-shot learners are plagued by referential ambiguityfrom distractors, and referring techniques are limited by a lack of visionpriors. We revisit Few-Shot Referring Video Object Segmentation (FSR-VOS) under a streamlined and practical formulation that segments tar-get objects using only query-level textual descriptions and a small set ofannotated support images. This dual conditioning enables coarse-to-finedisambiguation and precise boundary refinement with minimal annota-tion cost. To address the challenge of complex multi-source informationfusion, we propose the QST-SAM framework. Specifically, a SupportInstruction Compressor is introduced to refine and condense the sup-port information into compact representations. The Q-S-T Transformerfurther facilitates comprehensive integration of multi-modal supervisorycues from query, support, and text branches. The Multi-Priors Prompt-ing module provides the SAM2 prompt encoder with enriched guidancefor precise, context-aware segmentation. We further establish a new FS-RVOS benchmark derived from Ref-Youtube-VOS and Ref-DAVIS17, onwhich QST-SAM achieves state-of-the-art performance, validating its ro-bustness and effectiveness in segmentation.