2023
Learning to love diligent trolls: Accounting for rater effects in the dialogue safety task
EMNLP 2023short findings
Chatbots have the risk of generating offensive utterances, which must be avoided. Post-deployment, one way for a chatbot to continuously improve is to source utterance/label pairs from feedback by live users. However, among users are trolls, who provide training examples with incorrect labels. To de…