← Search

Yuanjie Chen

2 accepted papers

2026

Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models

ICLR 2026poster

While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become an increasingly important research focus. Despite their improvements, current MLLMs struggle significantly with high-le…

Cited by 0SourceScholar
2025

The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMs

ICLR 2025poster

Multi-modal Large Language Models (MLLMs) have exhibited impressive capability. However, recently many deficiencies of MLLMs have been found compared to human intelligence, $\textit{e.g.}$, hallucination. To drive the MLLMs study, the community dedicated efforts to building larger benchmarks with co…

Cited by 0SourcePDFScholar