2026
Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
AAAI 2026technical
Large language models (LLMs) have revolutionized AI applications, yet their high computational and memory demands hinder their widespread deployment. Existing compression techniques focus on intra-block optimizations (e.g., low-rank approximation or attention head pruning), while the repetitive laye