2025
Task-Informed Anti-Curriculum by Masking Improves Downstream Performance on Text
ACL 2025finding
Masked language modeling has become a widely adopted unsupervised technique to pre-train large language models (LLMs). However, the process of selecting tokens for masking is random, and the percentage of masked tokens is typically fixed for the entire training process. In this paper, we propose to…