← Search

Mateusz Klimaszewski

6 accepted papers

2025

An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)

ACL 2025long

Training state-of-the-art large language models requires vast amounts of clean and diverse textual data. However, building suitable multilingual datasets remains a challenge. In this work, we present HPLT v2, a collection of high-quality multilingual monolingual and parallel corpora, extending prior…

2025

EuroGEST: Investigating gender stereotypes in multilingual language models

EMNLP 2025

Large language models increasingly support multiple languages, yet most benchmarks for gender bias remain English-centric. We introduce EuroGEST, a dataset designed to measure gender-stereotypical reasoning in LLMs across English and 29 European languages. EuroGEST builds on an existing expert-infor

Cited by 0SourcePDFScholar
2025

Multilingual Data Filtering using Synthetic Data from Large Language Models

EMNLP 2025

Filtering data, particularly data scraped from the internet, has long been recognised as a means to improve model performance. Recent studies have shown that effective filters can be created by utilising Large Language Models (LLMs) to synthetically label data, which is then used to train smaller ne

Cited by 0SourcePDFScholar
2025

No Train but Gain: Language Arithmetic for training-free Language Adapters enhancement

COLING 2025main

Modular deep learning is the state-of-the-art solution for lifting the curse of multilinguality, preventing the impact of negative interference and enabling cross-lingual performance in Multilingual Pre-trained Language Models. However, a trade-off of this approach is the reduction in positive trans…

2024

Is Modularity Transferable? A Case Study through the Lens of Knowledge Distillation

COLING 2024main

The rise of Modular Deep Learning showcases its potential in various Natural Language Processing applications. Parameter-efficient fine-tuning (PEFT) modularity has been shown to work for various use cases, from domain adaptation to multilingual setups. However, all this work covers the case where t…

2021

COMBO: State-of-the-Art Morphosyntactic Analysis

EMNLP 2021system demonstrations

We introduce COMBO – a fully neural NLP system for accurate part-of-speech tagging, morphological analysis, lemmatisation, and (enhanced) dependency parsing. It predicts categorical morphosyntactic features whilst also exposes their vector representations, extracted from hidden layers. COMBO is an e…