← Search

Kebin Fang

1 accepted papers

2023

TLM: Token-Level Masking for Transformers

EMNLP 2023long main

Structured dropout approaches, such as attention dropout and DropHead, have been investigated to regularize the multi-head attention mechanism in Transformers. In this paper, we propose a new regularization scheme based on token-level rather than structure-level to reduce overfitting. Specifically,…

Cited by 0SourcecodeScholar