← Search

Maciej Piasecki

3 accepted papers

2024

BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language

COLING 2024main

The BEIR dataset is a large, heterogeneous benchmark for Information Retrieval (IR), garnering considerable attention within the research community. However, BEIR and analogous datasets are predominantly restricted to English language. Our objective is to establish extensive large-scale resources fo…

Cited by 9SourcePDFScholar
2024

Developing PUGG for Polish: A Modern Approach to KBQA, MRC, and IR Dataset Construction

ACL 2024findings

Advancements in AI and natural language processing have revolutionized machine-human language interactions, with question answering (QA) systems playing a pivotal role. The knowledge base question answering (KBQA) task, utilizing structured knowledge graphs (KG), allows for handling extensive knowle…

2022

This is the way: designing and compiling LEPISZCZE, a comprehensive NLP benchmark for Polish

NeurIPS 2022accept

The availability of compute and data to train larger and larger language models increases the demand for robust methods of benchmarking the true progress of LM training. Recent years witnessed significant progress in standardized benchmarking for English. Benchmarks such as GLUE, SuperGLUE, or KILT…