← Search

Baiyang Song

2 accepted papers

2026

Grounded Chain-of-Thought for Multimodal Large Language Models

CVPR 2026

Despite great progress, existing multimodal large language models (MLLMs) are still inferior in visual-spatial reasoning, which greatly impedes their trustworthy applications in scenarios such as Embodied AI. To facilitate the research, we propose a new MLLM task in this paper, called Grounded Chain

Cited by 0SourcecodeScholar
2026

KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs

AAAI 2026technical

Training-free video understanding methods leverage the strong image comprehension capabilities of pre-trained vision language models (VLMs) by treating videos as a sequences of static frames, thus obviating the need for costly video-specific training. However, this paradigm often suffers from severe

Cited by 0SourcePDFScholar