Contrast-CAT: Contrasting Activations for Enhanced Interpretability in Transformer-based Text Classifiers
Transformers have profoundly influenced AI research, but explaining their decisions remains challenging – even for relatively simpler tasks such as classification – which hinders trust and safe deployment in real-world applications. Although activation-based attribution methods effectively explain t