← Search

Zhenwei Shao

3 accepted papers

2026

VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding

CVPR 2026

Long-form video understanding remains challenging due to the extended temporal structure and dense multimodal cues. Despite recent progress, many existing approaches still rely on hand-crafted reasoning pipelines or employ token-consuming video preprocessing to guide MLLMs in autonomous reasoning. T

Cited by 0SourcecodeScholar
2025

Growing a Twig to Accelerate Large Vision-Language Models

ICCV 2025poster

Large vision-language models (VLMs) have demonstrated remarkable capabilities in open-world multimodal understanding, yet their high computational overheads pose great challenges for practical deployment. Some recent works have proposed methods to accelerate VLMs by pruning redundant visual tokens g…

2023

Prompting Large Language Models With Answer Heuristics for Knowledge-Based Visual Question Answering

CVPR 2023poster

Knowledge-based visual question answering (VQA) requires external knowledge beyond the image to answer the question. Early studies retrieve required knowledge from explicit knowledge bases (KBs), which often introduces irrelevant information to the question, hence restricting the performance of thei…