2025
Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences
EMNLP 2025
Current LLMs are trained to refuse potentially harmful input queries regardless of whether users actually had harmful intents, causing a tradeoff between safety and user experience. Through a study of 480 participants evaluating 3,840 query-response pairs, we examine how different refusal strategies