← Search

Xueyang Yu

3 accepted papers

2026

Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

CVPR 2026

Vision-language models (VLMs) excel at multimodal understanding, yet their text-only decoding forces them to verbalize visual reasoning, limiting performance on tasks that demand visual imagination. Recent attempts train VLMs to render explicit images, but the heavy image-generation pre-training oft

Cited by 0SourcecodeScholar
2025

TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation

NeurIPS 2025poster

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips totaling over 500 hours of high-quality 1080P human speech videos w…

Cited by 0SourceScholar
2025

VCA: Video Curious Agent for Long Video Understanding

ICCV 2025poster

Long video understanding poses unique challenges due to its temporal complexity and low information density. Recent works address this task by sampling numerous frames or incorporating auxiliary tools using LLMs, both of which result in high computational costs. In this work, we introduce a curiosit…

Cited by 0SourcePDFScholar