← Search

Tim Isbister

3 accepted papers

2025

SWEb: A Large Web Dataset for the Scandinavian Languages

ICLR 2025poster

This paper presents the hitherto largest pretraining dataset for the Scandinavian languages: the Scandinavian WEb (SWEb), comprising over one trillion tokens. The paper details the collection and processing pipeline, and introduces a novel model-based text extractor that significantly reduces comple…

Cited by 0SourcePDFScholar
2024

GPT-SW3: An Autoregressive Language Model for the Scandinavian Languages

COLING 2024main

This paper details the process of developing the first native large generative language model for the North Germanic languages, GPT-SW3. We cover all parts of the development process, from data collection and processing, training configuration and instruction finetuning, to evaluation, applications,…

2023

Superlim: A Swedish Language Understanding Evaluation Benchmark

EMNLP 2023long main

We present Superlim, a multi-task NLP benchmark and analysis platform for evaluating Swedish language models, a counterpart to the English-language (Super)GLUE suite. We describe the dataset, the tasks, the leaderboard and report the baseline results yielded by a reference implementation. The tested…

Cited by 0SourceScholar