← Search

Gilha lee

2 accepted papers

2026

SPLIT-VLM: Salience-Guided Partitioning towards Local Coverage for Importance-Aware Token Dropping in Vision-Language Models

ICML 2026poster

Large-scale vision–language models (VLMs) excel at multimodal reasoning, yet efficiency collapses when vision tokens—often orders of magnitude more than text—dominate compute and memory. Prior token-reduction strategies typically trade off salience (which is prone to position bias and incurs extra c…

Cited by 0SourceScholar
2026

WAVE: Window-Aware Vocabulary-Efficient Early-Exit for Training-Free LLM Acceleration

ICML 2026poster

Large language models (LLMs) incur substantial inference latency due to autoregressive decoding, in which each token requires a full forward pass through all transformer layers. Early-exit methods that terminate computation at intermediate layers offer a promising remedy, yet existing approaches suf…

Cited by 0SourceScholar