← Search

Anirudh Atmakuru

1 accepted papers

2025

Does quantization affect models’ performance on long-context tasks?

EMNLP 2025

Large language models (LLMs) now support context windows exceeding 128K tokens, but this comes with significant memory requirements and high inference latency. Quantization can mitigate these costs, but may degrade performance. In this work, we present the first systematic evaluation of quantized LL