← Search

Leyang Shen

3 accepted papers

2025

LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant

CVPR 2025poster

First-person video assistants are highly anticipated to enhance our daily life through online video dialogue. However, existing online video assistants often sacrifice assistant efficacy for real-time efficiency by processing low-frame-rate videos with coarse-grained visual features. To overcome the…

2024

LION: Empowering Multimodal Large Language Model with Dual-Level Visual Knowledge

CVPR 2024poster

Multimodal Large Language Models (MLLMs) have endowed LLMs with the ability to perceive and understand multi-modal signals. However most of the existing MLLMs mainly adopt vision encoders pretrained on coarsely aligned image-text pairs leading to insufficient extraction and reasoning of visual knowl…

2024

MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models

NeurIPS 2024poster

Multimodal large language models (MLLMs) have demonstrated impressive capabilities across various vision-language tasks. However, a generalist MLLM typically underperforms compared with a specialist MLLM on most VL tasks, which can be attributed to task interference. In this paper, we propose a mixt…