2025
Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision Models
NeurIPS 2025poster
Large multimodal models (LMMs) have gained impressive performance due to their outstanding capability in various understanding tasks. However, these models still suffer from some fundamental limitations related to robustness and generalization due to the alignment and correlation between visual and…