2026
SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token Pruning
AAAI 2026technical
Long-context inference for Large Language Models (LLMs) is heavily limited by high computational demands. While several existing methods optimize attention computation, they still process the full set of hidden states at each layer, limiting overall efficiency. In this work, we propose SlimInfer, an