On the Impact of Cross-Domain Data on German Language Models
Amin Dada, Aokun Chen, Cheng Peng, Kaleb E Smith, Ahmad Idrissi-Yaghir, Constantin Marc Seibold, Jianning Li, Lars Heiliger
Abstract
Traditionally, large language models have been either trained on general web crawls or domain-specific data. However, recent successes of generative large language models, have shed light on the benefits of cross-domain datasets. To examine the significance of prioritizing data diversity over quality, we present a German dataset comprising texts from five domains, along with another dataset aimed at containing high-quality data. Through training a series of models ranging between 122M and 750M parameters on both datasets, we conduct a comprehensive benchmark on multiple downstream tasks. Our findings demonstrate that the models trained on the cross-domain dataset outperform those trained on quality data alone, leading to improvements up to 4.45% over the previous state-of-the-art.
BibTeX
@inproceedings{
dada2023on,
title={On the Impact of Cross-Domain Data on German Language Models},
author={Amin Dada and Aokun Chen and Cheng Peng and Kaleb E Smith and Ahmad Idrissi-Yaghir and Constantin Marc Seibold and Jianning Li and Lars Heiliger and Christoph M. Friedrich and Daniel Truhn and Jan Egger and Jiang Bian and Jens Kleesiek and Yonghui Wu},
booktitle={The 2023 Conference on Empirical Methods in Natural Language Processing},
year={2023},
url={https://openreview.net/forum?id=8B9mL26NDT}
}