← Search

Desen Meng

3 accepted papers

2026

CaReBench: A Fine-grained Benchmark for Video Captioning and Retrieval

ICLR 2026poster

Video understanding, including video captioning and retrieval, is still a great challenge for video-language models (VLMs). The existing video retrieval and caption benchmarks only include short descriptions, limits their ability of detailed video understanding evaluation. To address this problem, w…

Cited by 0SourcecodeScholar
2025

LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization

NeurIPS 2025poster

We present LongVPO, a novel two‑stage Direct Preference Optimization framework that enables short‑context vision‑language models to robustly understand ultra‑long videos without any long‑video annotations. In Stage 1, we synthesize preference triples by anchoring questions to individual short clips,…

Cited by 0SourceScholar
2025

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay

ICCV 2025poster

Despite the remarkable performance of multimodal large language models (MLLMs) across diverse tasks, the substantial training and inference costs impede their advancement. In this paper, we propose p-MoD, an efficient MLLM architecture that significantly reduces training and inference costs while ma…