OpenAgentSafety: A Comprehensive Framework For Evaluating Real-World AI Agent Safety
Recent advances in AI agents capable of solving complex, everyday tasks-- from software engineering to customer service-- have enabled deployment in real-world settings, but their possibilities for unsafe behavior demands rigorous evaluation. While prior benchmarks have attempted to assess agent saf…