2025
Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
NeurIPS 2025poster
Transformer models have driven breakthroughs across various language tasks by their strong capability to learn rich contextual representations. Scaling them to improve representation, however, often demands substantial memory and compute costs, such as the Key-Value (KV) cache used during auto-regre…