← Search

Qihui Zhu

1 accepted papers

2026

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models

CVPR 2026

In multimodal large language models (MLLMs), the surge of visual tokens significantly increases the inference time and computational overhead, making them impractical for real-time or resource-constrained applications.Visual token pruning is a promising strategy for reducing the cost of MLLM inferen

Cited by 0SourcecodeScholar