2025
Are you sure? Measuring models bias in content moderation through uncertainty
EMNLP 2025
Automatic content moderation is crucial to ensuring safety in social media. Language Model-based classifiers are increasingly adopted for this task, but it has been shown that they perpetuate racial and social biases. Even if several resources and benchmark corpora have been developed to challenge t