← Search

Pierre-Carl Langlais

1 accepted papers

2026

Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training

ICLR 2026oral

Large Language Models (LLMs) are pre-trained on large data from different sources and domains. These data most often contain trillions of tokens with large portions of copyrighted or proprietary content, which hinders the usage of such models under AI legislation. This raises the need for truly open…

Cited by 0SourceScholar