2025
RealSafe: Quantifying Safety Risks of Language Agents in Real-World
COLING 2025main
We present RealSafe, an innovative evaluation framework that aims to rigorously assess the safety and reliability of large language model (LLM) agents in real application scenarios. RealSafe tracks the behavior of LLM agents in fourteen different application scenarios utilizing three contexts - stan…