← Search

Nairan Zhang

1 accepted papers

2025

VLM in a flash: I/O-Efficient Sparsification of Vision-Language Model via Neuron Chunking

NeurIPS 2025poster

Edge deployment of large Vision-Language Models (VLMs) increasingly relies on flash-based weight offloading, where activation sparsification is used to reduce I/O overhead. However, conventional sparsification remains model-centric, selecting neurons solely by activation magnitude and neglecting how…

Cited by 0SourceScholar