← Search

Bohan Zhuang*

2 accepted papers

2024

LongVLM: Efficient Long Video Understanding via Large Language Models

ECCV 2024oral

"Empowered by Large Language Models (LLMs), recent advancements in Video-based LLMs (VideoLLMs) have driven progress in various video understanding tasks. These models encode video representations through pooling or query aggregation over a vast number of visual tokens, making computational and memo…