2026
Learning to Lie: Adversarial Attacks on Human-AI Teams and LLMs
ICLR 2026poster
As artificial intelligence (AI) assistants become more widely adopted in safety-critical domains, it becomes important to develop safeguards against potential failures or adversarial attacks. A key prerequisite to developing these safeguards is understanding the ability of these AI assistants to mis…