2026
Unifying Adversarial Robustness and Training Across Text Scoring Models
ICML 2026poster
Research on adversarial robustness in language models is currently fragmented across applications and attacks, obscuring shared vulnerabilities. In this work, we propose unifying the study of adversarial robustness in text scoring models spanning dense retrievers, rerankers, and reward models. Unlik…