← Search

Merel Scholman

3 accepted papers

2024

DiscoGeM 2.0: A Parallel Corpus of English, German, French and Czech Implicit Discourse Relations

COLING 2024main

We present DiscoGeM 2.0, a crowdsourced, parallel corpus of 12,834 implicit discourse relations, with English, German, French and Czech data. We propose and validate a new single-step crowdsourcing annotation method and apply it to collect new annotations in German, French and Czech. The corpus was…

2024

Modeling Orthographic Variation Improves NLP Performance for Nigerian Pidgin

COLING 2024main

Nigerian Pidgin is an English-derived contact language and is traditionally an oral language, spoken by approximately 100 million people. No orthographic standard has yet been adopted, and thus the few available Pidgin datasets that exist are characterised by noise in the form of orthographic variat…

Cited by 2SourcePDFScholar
2022

Establishing Annotation Quality in Multi-label Annotations

COLING 2022main

In many linguistic fields requiring annotated data, multiple interpretations of a single item are possible. Multi-label annotations more accurately reflect this possibility. However, allowing for multi-label annotations also affects the chance that two coders agree with each other. Calculating inter…

Cited by 19SourcePDFScholar