← Search

Vanja Mladen Karan

4 accepted papers

2024

Denoising Labeled Data for Comment Moderation Using Active Learning

COLING 2024main

Noisily labeled textual data is ample on internet platforms that allow user-created content. Training models, such as offensive language detection models for comment moderation, on such data may prove difficult as the noise in the labels prevents the model to converge. In this work, we propose to us…

Cited by 1SourcePDFScholar
2023

LEDA: a Large-Organization Email-Based Decision-Dialogue-Act Analysis Dataset

ACL 2023findings

Collaboration increasingly happens online. This is especially true for large groups working on global tasks, with collaborators all around the globe. The size and distributed nature of such groups makes decision-making challenging. This paper proposes a set of dialog acts for the study of decision-m…

Cited by 3SourcePDFScholar
2023

Tracing Linguistic Markers of Influence in a Large Online Organisation

ACL 2023short

Social science and psycholinguistic research have shown that power and status affect how people use language in a range of domains. Here, we investigate a similar question in a large, distributed, consensus-driven community with little traditional power hierarchy – the Internet Engineering Task Forc…

Cited by 3SourcePDFScholar
2020

XHate-999: Analyzing and Detecting Abusive Language Across Domains and Languages

COLING 2020main

We present XHate-999, a multi-domain and multilingual evaluation data set for abusive language detection. By aligning test instances across six typologically diverse languages, XHate-999 for the first time allows for disentanglement of the domain transfer and language transfer effects in abusive lan…