2024
ChatGPT Doesn’t Trust Chargers Fans: Guardrail Sensitivity in Context
EMNLP 2024main
While the biases of language models in production are extensively documented, the biases of their guardrails have been neglected. This paper studies how contextual information about the user influences the likelihood of an LLM to refuse to execute a request. By generating user biographies that offer…