← Search

Yukun Feng

3 accepted papers

2022

Automatic Document Selection for Efficient Encoder Pretraining

EMNLP 2022main

Building pretrained language models is considered expensive and data-intensive, but must we increase dataset size to achieve better performance? We propose an alternative to larger training sets by automatically identifying smaller yet domain-representative subsets. We extend Cynical Data Selection,…

2022

Learn To Remember: Transformer with Recurrent Memory for Document-Level Machine Translation

NAACL 2022findings

The Transformer architecture has led to significant gains in machine translation. However, most studies focus on only sentence-level translation without considering the context dependency within documents, leading to the inadequacy of document-level coherence. Some recent research tried to mitigate…

Cited by 20SourcePDFScholar