← Search

Chaoyu Li

2 accepted papers

2026

FrameOracle: Learning What to See and How Much to See in Videos

ICML 2026poster

Vision-language models (VLMs) advance video understanding but operate under tight computational budgets, making performance dependent on selecting a small, high-quality subset of frames. Existing frame sampling strategies, such as uniform or fixed-budget selection, fail to adapt to variations in con…

Cited by 3SourceScholar
2025

VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding

CVPR 2025poster

Multimodal large language models (MLLMs) have recently shown significant advancements in video understanding, excelling in content reasoning and instruction-following tasks. However, hallucination, where models generate inaccurate or misleading content, remains underexplored in the video domain. Bui…