← Search

Chi-Hao Wu

2 accepted papers

2026

MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding

ICLR 2026poster

The rapid progress of large language models (LLMs) has laid the foundation for multimodal models. Nevertheless, visual language models (VLMs) still face significant computational overhead when scaled from images to the video domain. When video data is too large (due to high frame rates and long dura…

Cited by 0SourceScholar
2020

Fashion Captioning: Towards Generating Accurate Descriptions with Semantic Rewards

ECCV 2020poster

Generating accurate descriptions for online fashion items is important not only for enhancing customers' shopping experiences, but also for the increase of online sales. Besides the need of correctly presenting the attributes of items, the expressions in an enchanting style could better attract cust…