2024
Griffon: Spelling out All Object Locations at Any Granularity with Large Language Models
ECCV 2024poster
"Replicating the innate human ability to detect all objects based on free-form texts at any granularity remains a formidable challenge for Large Vision Language Models (LVLMs). Current LVLMs are predominantly constrained to locate a single, pre-existing object. This limitation leads to a compromise…