← Search

Qiankun Ma

2 accepted papers

2026

ApET: Approximation-Error Guided Token Compression for Efficient VLMs

CVPR 2026

Recent Vision-Language Models (VLMs) have demonstrated remarkable multimodal understanding capabilities, yet the redundant visual tokens incur prohibitive computational overhead and degrade inference efficiency. Prior studies typically relies on [CLS] attention or text-vision cross-attention to iden

Cited by 0SourcecodeScholar
2023

Rethinking Safe Semi-supervised Learning: Transferring the Open-set Problem to A Close-set One

ICCV 2023poster

Conventional semi-supervised learning (SSL) lies in the close-set assumption that the labeled and unlabeled sets contain data with the same seen classes, called in-distribution (ID) data. In contrast, safe SSL investigates a more challenging open-set problem where unlabeled set may involve some out-…

Cited by 12PDFScholar