← Search

Phillip Y. Lee

3 accepted papers

2026

Token Warping Helps MLLMs Look from Nearby Viewpoints

CVPR 2026

Can warping tokens, rather than pixels, help multimodal large language models (MLLMs) understand how a scene appears from a nearby viewpoint? While MLLMs perform well on visual reasoning, they remain fragile to viewpoint changes, as pixel-wise warping is highly sensitive to small depth errors and of

Cited by 0SourcecodeScholar
2025

Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation

ICCV 2025poster

We present a framework for perspective-aware reasoning in vision-language models (VLMs) through mental imagery simulation. Perspective-taking, the ability to perceive an environment or situation from an alternative viewpoint, is a key benchmark for human-level visual understanding, essential for env…

Cited by 0SourcePDFScholar