2025
AgentAuditor: Human-level Safety and Security Evaluation for LLM Agents
NeurIPS 2025poster
Despite the rapid advancement of LLM-based agents, the reliable evaluation of their safety and security remains a significant challenge. Existing rule-based or LLM-based evaluators often miss dangers in agents' step-by-step actions, overlook subtle meanings, fail to see how small issues compound, an…