← Search

Victoria R Li

1 accepted papers

2024

ChatGPT Doesn’t Trust Chargers Fans: Guardrail Sensitivity in Context

EMNLP 2024main

While the biases of language models in production are extensively documented, the biases of their guardrails have been neglected. This paper studies how contextual information about the user influences the likelihood of an LLM to refuse to execute a request. By generating user biographies that offer…