2024
Debiasing Text Safety Classifiers through a Fairness-Aware Ensemble
EMNLP 2024industry
Increasing use of large language models (LLMs) demand performant guardrails to ensure the safety of inputs and outputs of LLMs. When these safeguards are trained on imbalanced data, they can learn the societal biases. We present a light-weight, post-processing method for mitigating counterfactual fa…