2025
MULTIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and Modalities
EMNLP 2025
The emerging capabilities of large language models (LLMs) have sparked concerns about their immediate potential for harmful misuse. The core approach to mitigate these concerns is the detection of harmful queries to the model. Current detection approaches are fallible, and are particularly susceptib