2024
A Novel Metric for Measuring the Robustness of Large Language Models in Non-adversarial Scenarios
EMNLP 2024finding
We evaluate the robustness of several large language models on multiple datasets. Robustness here refers to the relative insensitivity of the model’s answers to meaning-preserving variants of their input. Benchmark datasets are constructed by introducing naturally-occurring, non-malicious perturbati…