← Search

Brian Formento

2 accepted papers

2025

Confidence Elicitation: A New Attack Vector for Large Language Models

ICLR 2025poster

A fundamental issue in deep learning has been adversarial robustness. As these systems have scaled, such issues have persisted. Currently, large language models (LLMs) with billions of parameters suffer from adversarial attacks just like their earlier, smaller counterparts. However, the threat model…

2024

SemRoDe: Macro Adversarial Training to Learn Representations that are Robust to Word-Level Attacks

NAACL 2024long

Language models (LMs) are indispensable tools for natural language processing tasks, but their vulnerability to adversarial attacks remains a concern. While current research has explored adversarial training techniques, their improvements to defend against word-level attacks have been limited. In th…