2023
CFL: Causally Fair Language Models Through Token-level Attribute Controlled Generation
ACL 2023findings
We propose a method to control the attributes of Language Models (LMs) for the text generation task using Causal Average Treatment Effect (ATE) scores and counterfactual augmentation. We explore this method, in the context of LM detoxification, and propose the Causally Fair Language (CFL) architectu…