← Search

Kristian Lum

4 accepted papers

2025

Bias in Language Models: Beyond Trick Tests and Towards RUTEd Evaluation

ACL 2025long

Standard bias benchmarks used for large language models (LLMs) measure the association between social attributes in model inputs and single-word model outputs. We test whether these benchmarks are robust to lengthening the model outputs via a more realistic user prompt, in the commonly studied domai…

Cited by 0SourcePDFScholar
2024

STAR: SocioTechnical Approach to Red Teaming Language Models

EMNLP 2024main

This research introduces STAR, a sociotechnical framework that improves on current best practices for red teaming safety of large language models. STAR makes two key contributions: it enhances steerability by generating parameterised instructions for human red teamers, leading to improved coverage o…

Cited by 13SourcePDFScholar
2021

It's COMPASlicated: The Messy Relationship between RAI Datasets and Algorithmic Fairness Benchmarks

NeurIPS 2021poster

Risk assessment instrument (RAI) datasets, particularly ProPublica’s COMPAS dataset, are commonly used in algorithmic fairness papers due to benchmarking practices of comparing algorithms on datasets used in prior work. In many cases, this data is used as a benchmark to demonstrate good performance…

Cited by 125SourceScholar