2024
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction
ACL 2024long
Social biases in LLMs are usually measured via bias benchmark datasets. Current benchmarks have limitations in scope, grounding, quality, and human effort required. Previous work has shown success with a community-sourced, rather than crowd-sourced, approach to benchmark development. However, this w…