← Search

Yuxuan Ge

2 accepted papers

2026

Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs

ICLR 2026poster

Multimodal Large Language Models (MLLMs) combine the linguistic strengths of LLMs with the ability to process multimodal data, enabling them to address a broader range of tasks. This progression highlights a shift from language-only reasoning to integrated vision–language reasoning in children's dev…

Cited by 0SourcecodeScholar
2024

Tri-Modal Motion Retrieval by Learning a Joint Embedding Space

CVPR 2024highlight

Text-to-motion tasks have been the focus of recent advancements in the human motion domain. However the performance of text-to-motion tasks have not reached its potential primarily due to the lack of motion datasets and the pronounced gap between the text and motion modalities. To mitigate this chal…

Cited by 5SourcePDFScholar