2024
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling
NeurIPS 2024poster
Constrained by the separate encoding of vision and language, existing grounding and referring segmentation works heavily rely on bulky Transformer-based fusion en-/decoders and a variety of early-stage interaction technologies. Simultaneously, the current mask visual language modeling (MVLM) fails t…