← Search

Jiacheng Luo

1 accepted papers

2026

HumorReject: Decoupling LLM Safety from Refusal Prefix via a Little Humor

AAAI 2026technical

Large Language Models (LLMs) commonly rely on explicit refusal prefixes for safety, making them vulnerable to prefix injection attacks. We introduce HumorReject, a novel data-driven approach that reimagines LLM safety by decoupling it from refusal prefixes through humor as an indirect refusal strate

Cited by 0SourcePDFScholar