2025
The Unreasonable Ineffectiveness of the Deeper Layers
ICLR 2025poster
How is knowledge stored in an LLM’s weights? We study this via layer pruning: if removing a certain layer does not affect model performance in common question-answering benchmarks, then the weights in that layer are not necessary for storing the knowledge needed to answer those questions. To find th…