2026
Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection
AAAI 2026technical
The opaque nature of deep learning models presents significant challenges for the ethical deployment of hate speech detection systems. To address this limitation, we introduce Supervised Rational Attention (SRA), a framework that explicitly aligns model attention with human rationales, improving bot