2026
Escaping the Likelihood Trap: Geometric Diversity Optimization for Long-Form Image Captioning
ICML 2026poster
The utility of Vision-Language Models (VLMs) in reasoning and auditing tasks hinges on their ability to exhaustively describe visual scenes. However, current models exhibit a pathology we term the Likelihood Trap: standard alignment objectives, specifically MLE and KL-regularization, drive generatio…