2025
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
ACL 2025long
While long-context inference is crucial for advancing large language model (LLM) applications, its prefill speed remains a significant bottleneck. Current approaches, including sequence parallelism strategies and compute reduction through approximate attention mechanisms, still fall short of deliver…