← Search

Fredrik Carlsson

6 accepted papers

2026

SEDRAS: Symbolically Evaluated Deep Research And Science

ICML 2026poster

As the reasoning capabilities of Large Language Models (LLMs) expand, evaluating true inductive generalization on entirely unseen data becomes increasingly challenging. To this end, we introduce a modular in-context learning evaluation framework, that is scalable and extendable across its separate m…

Cited by 0SourceScholar
2025

The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation

ICLR 2025poster

This paper introduces the counter-intuitive generalization results of overfitting pre-trained large language models (LLMs) on very small datasets. In the setting of open-ended text generation, it is well-documented that LLMs tend to generate repetitive and dull sequences, a phenomenon that is especi…

Cited by 1SourcePDFScholar
2024

Branch-GAN: Improving Text Generation with (not so) Large Language Models

ICLR 2024poster

The current advancements in open domain text generation have been spearheaded by Transformer-based large language models. Leveraging efficient parallelization and vast training datasets, these models achieve unparalleled text generation capabilities. Even so, current models are known to suffer from…

Cited by 3SourcePDFScholar
2024

GPT-SW3: An Autoregressive Language Model for the Scandinavian Languages

COLING 2024main

This paper details the process of developing the first native large generative language model for the North Germanic languages, GPT-SW3. We cover all parts of the development process, from data collection and processing, training configuration and instruction finetuning, to evaluation, applications,…

2022

Fine-Grained Controllable Text Generation Using Non-Residual Prompting

ACL 2022long

The introduction of immensely large Causal Language Models (CLMs) has rejuvenated the interest in open-ended text generation. However, controlling the generative process for these Transformer-based models is at large an unsolved problem. Earlier work has explored either plug-and-play decoding strate…

2021

Semantic Re-tuning with Contrastive Tension

ICLR 2021poster

Extracting semantically useful natural language sentence representations from pre-trained deep neural networks such as Transformers remains a challenge. We first demonstrate that pre-training objectives impose a significant task bias onto the final layers of models with a layer-wise survey of the Se…

Cited by 98SourcePDFScholar