← Search

Tobias Norlund

3 accepted papers

2025

SWEb: A Large Web Dataset for the Scandinavian Languages

ICLR 2025poster

This paper presents the hitherto largest pretraining dataset for the Scandinavian languages: the Scandinavian WEb (SWEb), comprising over one trillion tokens. The paper details the collection and processing pipeline, and introduces a novel model-based text extractor that significantly reduces comple…

Cited by 0SourcePDFScholar
2023

Surface-Based Retrieval Reduces Perplexity of Retrieval-Augmented Language Models

ACL 2023short

Augmenting language models with a retrieval mechanism has been shown to significantly improve their performance while keeping the number of parameters low. Retrieval-augmented models commonly rely on a semantic retrieval mechanism based on the similarity between dense representations of the query ch…

2023

The Effect of Scaling, Retrieval Augmentation and Form on the Factual Consistency of Language Models

EMNLP 2023long main

Large Language Models (LLMs) make natural interfaces to factual knowledge, but their usefulness is limited by their tendency to deliver inconsistent answers to semantically equivalent questions. For example, a model might supply the answer "Edinburgh" to "Anne Redpath passed away in X." and "London"…

Cited by 0SourcecodeScholar