2024
SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant
ECCV 2024poster
"Recent advances in vision-language models have shown notable generalization in broad tasks through visual instruction tuning. However, bridging the gap between the pre-trained vision encoder and the large language models (LLMs) becomes the whole network’s bottleneck. To improve cross-modality align…