2023
Multi-armed bandits for resource efficient, online optimization of language model pre-training: the use case of dynamic masking
ACL 2023findings
We design and evaluate a Bayesian optimization framework for resource efficient pre-training of Transformer-based language models (TLMs). TLM pre-training requires high computational resources and introduces many unresolved design choices, such as selecting its pre-training hyperparameters.We propos…