2026
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding
ICLR 2026poster
We introduce FLARE, a family of vision language models (VLMs) with a fully vision-language alignment and integration paradigm. Unlike existing approaches that rely on single MLP projectors for modality alignment and defer cross-modal interaction to LLM decoding, FLARE achieves deep, dynamic integrat…