2025
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
EMNLP 2025
Large Language Models (LLMs) are highly sensitive to subtle, non-semantic variations in prompt phrasing and formatting. In this work, we present the first systematic evaluation of 4 methods for improving prompt robustness within a unified experimental framework. We benchmark these techniques on 8 mo