2024
BeanCounter: A low-toxicity, large-scale, and open dataset of business-oriented text
NeurIPS 2024poster
Many of the recent breakthroughs in language modeling have resulted from scaling effectively the same model architecture to larger datasets. In this vein, recent work has highlighted performance gains from increasing training dataset size and quality, suggesting a need for novel sources of large-sca…