← Search

Røskva Bjørgfinsdóttir

1 accepted papers

2026

Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection

AAAI 2026technical

The opaque nature of deep learning models presents significant challenges for the ethical deployment of hate speech detection systems. To address this limitation, we introduce Supervised Rational Attention (SRA), a framework that explicitly aligns model attention with human rationales, improving bot

Cited by 0SourcePDFScholar