← Search

Nish Sinnadurai

2 accepted papers

2026

REAP the Experts: Why Pruning Prevails for One-Shot MoE compression

ICLR 2026poster

Sparsely-activated Mixture-of-Experts (SMoE) models offer efficient pre-training and low latency but their large parameter counts create significant memory overhead, motivating research into expert compression. Contrary to recent findings favouring expert *merging* on discriminative benchmarks, we f…

Cited by 0SourcecodeScholar
2025

MASSV: Multimodal Adaptation and Self-Data Distillation for Speculative Decoding of Vision-Language Models

EMNLP 2025

Speculative decoding significantly accelerates language model inference by enabling a lightweight draft model to propose multiple tokens that a larger target model verifies simultaneously. However, applying this technique to vision-language models (VLMs) presents two fundamental challenges: small la

Cited by 0SourcePDFScholar