2026
HumorReject: Decoupling LLM Safety from Refusal Prefix via a Little Humor
AAAI 2026technical
Large Language Models (LLMs) commonly rely on explicit refusal prefixes for safety, making them vulnerable to prefix injection attacks. We introduce HumorReject, a novel data-driven approach that reimagines LLM safety by decoupling it from refusal prefixes through humor as an indirect refusal strate