← Search

Hokin Deng

3 accepted papers

2026

Vision Language Models Cannot Reason About Physical Transformation

ICML 2026poster

Understanding physical transformations is fundamental for reasoning in dynamic environments. While Vision Language Models (VLMs) show promise in embodied applications, whether they genuinely understand physical transformations remains unclear. We introduce ***ConservationBench*** evaluating ***conse…

Cited by 0SourceScholar
2025

Core Knowledge Deficits in Multi-Modal Language Models

ICML 2025poster

While Multi-modal Large Language Models (MLLMs) demonstrate impressive abilities over high-level perception and reasoning, their robustness in the wild remains limited, often falling short on tasks that are intuitive and effortless for humans. We examine the hypothesis that these deficiencies stem f…

Cited by 0SourcePDFScholar