2025
TAI3: Testing Agent Integrity in Interpreting User Intent
NeurIPS 2025poster
LLM agents are increasingly deployed to automate real-world tasks by invoking APIs through natural language instructions. While powerful, they often suffer from misinterpretation of user intent, leading to the agent’s actions that diverge from the user’s intended goal, especially as external toolkit…