← Search

Daniel Hesslow

3 accepted papers

2023

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data Only

NeurIPS 2023poster

Large language models are commonly trained on a mixture of filtered web data and curated ``high-quality'' corpora, such as social media conversations, books, or technical papers. This curation process is believed to be necessary to produce performant models with broad zero-shot generalization abilit…

Cited by 147SourcePDFScholar
2022

What Language Model Architecture and Pretraining Objective Works Best for Zero-Shot Generalization?

ICML 2022spotlight

Large pretrained Transformer language models have been shown to exhibit zero-shot generalization, i.e. they can perform a wide variety of tasks that they were not explicitly trained on. However, the architectures and pretraining objectives used across state-of-the-art models differ significantly, an…

2022

What Language Model to Train if You Have One Million GPU Hours?

EMNLP 2022finding

The crystallization of modeling methods around the Transformer architecture has been a boon for practitioners. Simple, well-motivated architectural variations can transfer across tasks and scale, increasing the impact of modeling research. However, with the emergence of state-of-the-art 100B+ parame…