← Search

Joakim Nivre

7 accepted papers

2026

SEDRAS: Symbolically Evaluated Deep Research And Science

ICML 2026poster

As the reasoning capabilities of Large Language Models (LLMs) expand, evaluating true inductive generalization on entirely unseen data becomes increasingly challenging. To this end, we introduce a modular in-context learning evaluation framework, that is scalable and extendable across its separate m…

Cited by 0SourceScholar
2025

The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation

ICLR 2025poster

This paper introduces the counter-intuitive generalization results of overfitting pre-trained large language models (LLMs) on very small datasets. In the setting of open-ended text generation, it is well-documented that LLMs tend to generate repetitive and dull sequences, a phenomenon that is especi…

Cited by 1SourcePDFScholar
2024

Branch-GAN: Improving Text Generation with (not so) Large Language Models

ICLR 2024poster

The current advancements in open domain text generation have been spearheaded by Transformer-based large language models. Leveraging efficient parallelization and vast training datasets, these models achieve unparalleled text generation capabilities. Even so, current models are known to suffer from…

Cited by 3SourcePDFScholar
2024

UCxn: Typologically Informed Annotation of Constructions Atop Universal Dependencies

COLING 2024main

The Universal Dependencies (UD) project has created an invaluable collection of treebanks with contributions in over 140 languages. However, the UD annotations do not tell the full story. Grammatical constructions that convey meaning through a particular combination of several morphosyntactic elemen…

2022

Fine-Grained Controllable Text Generation Using Non-Residual Prompting

ACL 2022long

The introduction of immensely large Causal Language Models (CLMs) has rejuvenated the interest in open-ended text generation. However, controlling the generative process for these Transformer-based models is at large an unsolved problem. Earlier work has explored either plug-and-play decoding strate…

2020

Understanding Pure Character-Based Neural Machine Translation: The Case of Translating Finnish into English

COLING 2020main

Recent work has shown that deeper character-based neural machine translation (NMT) models can outperform subword-based models. However, it is still unclear what makes deeper character-based models successful. In this paper, we conduct an investigation into pure character-based models in the case of…

Cited by 9SourcePDFScholar