← Search

Susav Shrestha

1 accepted papers

2025

Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity

NeurIPS 2025poster

Accelerating large language model (LLM) inference is critical for real-world deployments requiring high throughput and low latency. Contextual sparsity, where each token dynamically activates only a small subset of the model parameters, shows promise but does not scale to large batch sizes due to un…

Cited by 0SourcecodeScholar