LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
Recent advancements in Large Language Models (LLMs) have spurred interest in numerous applications requiring robust long-range capabilities, essential for processing extensive input contexts and continuously generating extended outputs. As sequence lengths increase, the number of Key-Value (KV) pair…