← Search

Magnus Sahlgren

7 accepted papers

2026

Position: If open source is to win, it must go public

ICML 2026spotlight

Open source projects have made incredible progress in producing widely usable machine learning models and systems, but open source alone will face challenges in fully democratizing access to AI. Unlike previous generations of open source software, open source and open weight AI models require substa…

Cited by 0SourceScholar
2025

SWEb: A Large Web Dataset for the Scandinavian Languages

ICLR 2025poster

This paper presents the hitherto largest pretraining dataset for the Scandinavian languages: the Scandinavian WEb (SWEb), comprising over one trillion tokens. The paper details the collection and processing pipeline, and introduces a novel model-based text extractor that significantly reduces comple…

Cited by 0SourcePDFScholar
2024

Branch-GAN: Improving Text Generation with (not so) Large Language Models

ICLR 2024poster

The current advancements in open domain text generation have been spearheaded by Transformer-based large language models. Leveraging efficient parallelization and vast training datasets, these models achieve unparalleled text generation capabilities. Even so, current models are known to suffer from…

Cited by 3SourcePDFScholar
2024

GPT-SW3: An Autoregressive Language Model for the Scandinavian Languages

COLING 2024main

This paper details the process of developing the first native large generative language model for the North Germanic languages, GPT-SW3. We cover all parts of the development process, from data collection and processing, training configuration and instruction finetuning, to evaluation, applications,…

2023

Superlim: A Swedish Language Understanding Evaluation Benchmark

EMNLP 2023long main

We present Superlim, a multi-task NLP benchmark and analysis platform for evaluating Swedish language models, a counterpart to the English-language (Super)GLUE suite. We describe the dataset, the tasks, the leaderboard and report the baseline results yielded by a reference implementation. The tested…

Cited by 0SourceScholar
2022

Fine-Grained Controllable Text Generation Using Non-Residual Prompting

ACL 2022long

The introduction of immensely large Causal Language Models (CLMs) has rejuvenated the interest in open-ended text generation. However, controlling the generative process for these Transformer-based models is at large an unsolved problem. Earlier work has explored either plug-and-play decoding strate…

2021

Semantic Re-tuning with Contrastive Tension

ICLR 2021poster

Extracting semantically useful natural language sentence representations from pre-trained deep neural networks such as Transformers remains a challenge. We first demonstrate that pre-training objectives impose a significant task bias onto the final layers of models with a layer-wise survey of the Se…

Cited by 98SourcePDFScholar