← Search

Mark Endo

4 accepted papers

2026

Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusions

CVPR 2026

Large Vision-Language Models (VLMs) often answer classic visual illusions "correctly" on original images, yet persist with the same responses when illusion factors are inverted, even though the visual change is obvious to humans. This raises a fundamental question: do VLMs perceive visual changes or

Cited by 0SourceScholar
2026

Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models

CVPR 2026

Scaling up multimodal models has enabled remarkable advances in visual understanding and reasoning, but practical demands call for smaller, efficient systems. In this work, we conduct a principled analysis of downscaling intelligence in multimodal models, examining how reduced large language model (

Cited by 0SourcecodeScholar
2025

Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration

ICCV 2025poster

Recent works on accelerating Vision-Language Models achieve strong performance across a variety of vision-language tasks despite highly compressing visual information. In this work, we examine the popular acceleration approach of early pruning of visual tokens inside the language model. Surprisingly…

Cited by 0SourcePDFScholar