2026
Multi-Feature Quantized Self-Attention for Fair Large Language Models
ICLR 2026poster
Large language models (LLMs) often encode social biases tied to sensitive features such as race and gender, undermining fairness in downstream tasks even after instruction tuning. Conventional debiasing methods require expensive fine-tuning, are tied to specific architectures, or operate only at the…