2024
Jump to Conclusions: Short-Cutting Transformers with Linear Transformations
COLING 2024main
Transformer-based language models create hidden representations of their inputs at every layer, but only use final-layer representations for prediction. This obscures the internal decision-making process of the model and the utility of its intermediate representations. One way to elucidate this is t…