2026
RSAgent: Learning to Reason and Act via Multi-Turn Tool Invocations for Text-Guided Segmentation
ICML 2026poster
Text-guided object segmentation requires both cross-modal reasoning and pixel grounding abilities. Most recent methods treat it as a single forward pass, where the model directly predicts pixel prompts to a segmentation model, which limits verification, refocusing and refinement when initial localiz…