2025
Confidence Elicitation: A New Attack Vector for Large Language Models
ICLR 2025poster
A fundamental issue in deep learning has been adversarial robustness. As these systems have scaled, such issues have persisted. Currently, large language models (LLMs) with billions of parameters suffer from adversarial attacks just like their earlier, smaller counterparts. However, the threat model…