← Search

Arjun Bhagoji

2 accepted papers

2025

Adapting to Evolving Adversaries with Regularized Continual Robust Training

ICML 2025poster

Robust training methods typically defend against specific attack types, such as $\ell_p$ attacks with fixed budgets, and rarely account for the fact that defenders may encounter new attacks over time. A natural solution is to adapt the defended model to new adversaries as they arise via fine-tuning…

2025

Silencing Empowerment, Allowing Bigotry: Auditing the Moderation of Hate Speech on Twitch

ACL 2025long

To meet the demands of content moderation, online platforms have resorted to automated systems. Newer forms of real-time engagement (e.g., users commenting on live streams) on platforms like Twitch exert additional pressures on the latency expected of such moderation systems. Despite their prevalenc…