← Search

Chenxi Liao

2 accepted papers

2026

IF-VidCap: Can Video Caption Models Follow Instructions?

ICLR 2026poster

Although Multimodal Large Language Models (MLLMs) have demonstrated proficiency in video captioning, practical applications require captions that follow specific user instructions rather than generating exhaustive, unconstrained descriptions. Current benchmarks, however, primarily assess descriptiv…

Cited by 0SourcecodeScholar
2026

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

ICML 2026poster

Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped benchmarks that fail to capture cross-modal alignment, instruction…

Cited by 0SourceScholar