2024
Depth-Wise Attention (DWAtt): A Layer Fusion Method for Data-Efficient Classification
COLING 2024main
Language Models pretrained on large textual data have been shown to encode different types of knowledge simultaneously. Traditionally, only the features from the last layer are used when adapting to new tasks or data. We put forward that, when using or finetuning deep pretrained models, intermediate…