2025
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
NeurIPS 2025poster
Aligning visual features with language embeddings is a key challenge in vision-language models (VLMs). The performance of such models hinges on having a good connector that maps visual features generated by a vision encoder to a shared embedding space with the LLM while preserving semantic similarit…