2026
Sliding-Window Merging for Compacting Patch-Redundant Layers in LLMs
AAAI 2026technical
Depth-wise pruning accelerates LLM inference in resource-constrained scenarios but suffers from performance degradation due to indiscriminate removal of entire Transformer layers. This paper reveals ``Patch-Like