← Search

Shuo Gao

2 accepted papers

2026

IF-VidCap: Can Video Caption Models Follow Instructions?

ICLR 2026poster

Although Multimodal Large Language Models (MLLMs) have demonstrated proficiency in video captioning, practical applications require captions that follow specific user instructions rather than generating exhaustive, unconstrained descriptions. Current benchmarks, however, primarily assess descriptiv…

Cited by 0SourcecodeScholar
2025

Explicit Spatial Hint and Implicit Logits Relation: Distilling Heterogeneous Knowledge From Vision Transformer to CNN

ICASSP 2025accepted

A lightweight Convolutional Neural Network (CNN) typically requires knowledge transfer from a large powerful network before it is employed in resource-limited edge devices. Vision Transformer (ViT) possesses an unparalleled capability for global modeling but remains largely unexplored in Knowledge D…

Cited by 0SourceScholar