← Search

Tu Trinh

3 accepted papers

2025

Aligned LLMs Are Not Aligned Browser Agents

ICLR 2025poster

For safety reasons, large language models (LLMs) are trained to refuse harmful user instructions, such as assisting dangerous activities. We study an open question in this work: does the desired safety refusal, typically enforced in chat contexts, generalize to non-chat and agentic use cases? Unlike…

Cited by 0SourcePDFScholar
2024

A StrongREJECT for Empty Jailbreaks

NeurIPS 2024poster

Most jailbreak papers claim the jailbreaks they propose are highly effective, often boasting near-100% attack success rates. However, it is perhaps more common than not for jailbreak developers to substantially exaggerate the effectiveness of their jailbreaks. We suggest this problem arises because…

2022

Efficient Game-Theoretic Planning With Prediction Heuristic for Socially-Compliant Autonomous Driving

RA-L 2022

Planning under social interactions with other agents is an essential problem for autonomous driving. As the actions of the autonomous vehicle in the interactions affect and are also affected by other agents, autonomous vehicles need to efficiently infer the reaction of the other agents. Most existin

Cited by 23SourceScholar