2025
CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs
NeurIPS 2025poster
Large language models (LLMs) are increasingly deployed in medical contexts, raising critical concerns about safety, alignment, and susceptibility to adversarial manipulation. While prior benchmarks assess model refusal capabilities for harmful prompts, they often lack clinical specificity, graded ha…