← Search

Zeliang Zong

3 accepted papers

2026

Less Token, More Signal: MoE Expert Pruning via Critical Token Selection

ICML 2026poster

Mixture-of-Experts (MoE) architectures provide strong scalability for large language models, but their large expert parameter footprint poses challenges for efficient deployment. Expert pruning is widely used to reduce model size and inference cost; however, existing approaches are token-agnostic, t…

Cited by 0SourceScholar
2026

One Layer's Trash is Another Layer's Treasure: Adaptive Layer-wise Visual Token Selection in LVLMs

CVPR 2026

Large Vision-Language Models (LVLMs) have achieved remarkable success across diverse multimodal tasks, yet their practical deployment remains constrained by the computational burden arising from lengthy visual tokens. While visual token pruning has emerged as a promising solution, existing methods s

Cited by 0SourceScholar
2025

1+1>2: A Synergistic Sparse and Low-Rank Compression Method for Large Language Models

EMNLP 2025

Large Language Models (LLMs) have demonstrated remarkable proficiency in language comprehension and generation; however, their widespread adoption is constrained by substantial bandwidth and computational demands. While pruning and low-rank approximation have each demonstrated promising performance