RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference
DeepSeek-OCR leverages visual–text compression to reduce long-text processing costs and accelerate inference, yet visual tokens remain prone to redundant textual and structural information. Moreover, current token pruning methods for conventional vision–language models (VLMs) fail to preserve textua…