← Search

Arkadiusz Janz

3 accepted papers

2025

PLLuM-Align: Polish Preference Dataset for Large Language Model Alignment

EMNLP 2025

Alignment is the critical process of minimizing harmful outputs by teaching large language models (LLMs) to prefer safe, helpful and appropriate responses. While the majority of alignment research and datasets remain overwhelmingly English-centric, ensuring safety across diverse linguistic and cultu

Cited by 0SourcePDFScholar
2024

BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language

COLING 2024main

The BEIR dataset is a large, heterogeneous benchmark for Information Retrieval (IR), garnering considerable attention within the research community. However, BEIR and analogous datasets are predominantly restricted to English language. Our objective is to establish extensive large-scale resources fo…

Cited by 9SourcePDFScholar
2022

This is the way: designing and compiling LEPISZCZE, a comprehensive NLP benchmark for Polish

NeurIPS 2022accept

The availability of compute and data to train larger and larger language models increases the demand for robust methods of benchmarking the true progress of LM training. Recent years witnessed significant progress in standardized benchmarking for English. Benchmarks such as GLUE, SuperGLUE, or KILT…