← Search

Shanwei Zhao

2 accepted papers

2026

QuietPrune: Query-Guided Early Token Pruning for Vision-Language Models

CVPR 2026

Vision-language models (VLMs) demonstrate powerful capabilities in multimodal tasks. However, the large number of visual tokens imposes a significant computational cost. In this paper, we propose QuietPrune, a QUery-guIded Early Token Pruning method to remove redundant visual tokens within VLMs, the

Cited by 0SourcecodeScholar
2026

SPEED-Q: Staged Processing with Enhanced Distillation Towards Efficient Low-Bit On-Device VLM Quantization

AAAI 2026technical

Deploying Vision-Language Models (VLMs) on edge devices (e.g., smartphones and robots) is crucial for enabling low-latency and privacy-preserving intelligent applications. Given the resource constraints of these devices, quantization offers a promising solution by improving memory efficiency and red

Cited by 0SourcePDFScholar