2026
ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
AAAI 2026technical
Large Multimodal Models (LMMs) often face a modality representation gap during pretraining: while language embeddings remain stable, visual representations are highly sensitive to contextual noise (e.g., background clutter). To address this issue, we introduce a visual comprehension stage, which we