2025
Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach
NeurIPS 2025poster
Large language models (LLMs) typically generate identical or similar responses for all users given the same prompt, posing serious safety risks in high-stakes applications where user vulnerabilities differ widely. Existing safety evaluations primarily rely on context-independent metrics—such as fact…