← Search

Pieter Fivez

2 accepted papers

2025

In Benchmarks We Trust ... Or Not?

EMNLP 2025

Standardized benchmarks are central to evaluating and comparing model performance in Natural Language Processing (NLP). However, Large Language Models (LLMs) have exposed shortcomings in existing benchmarks, and so far there is no clear solution. In this paper, we survey a wide scope of benchmarking

Cited by 1SourcePDFScholar
2021

Mapping probability word problems to executable representations

EMNLP 2021main

While solving math word problems automatically has received considerable attention in the NLP community, few works have addressed probability word problems specifically. In this paper, we employ and analyse various neural models for answering such word problems. In a two-step approach, the problem t…

Cited by 12SourcePDFScholar