← Search

Tim Beyer

3 accepted papers

2026

A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness

ICML 2026poster

Automated \enquote{LLM-as-a-Judge} frameworks have become the de facto standard for scalable evaluation across natural language processing. For instance, in safety evaluation, these judges are relied upon to evaluate harmfulness in order to benchmark the robustness of safety against adversarial atta…

Cited by 0SourceScholar
2026

Position: LLM-Safety Evaluations Lack Robustness

ICML 2026poster

In this position paper, we argue that current safety alignment research efforts for large language models are hindered by many intertwined sources of noise, such as small datasets, methodological inconsistencies, and unreliable evaluation setups. This can, at times, make it impossible to evaluate an…

Cited by 0SourceScholar
2026

Sampling-aware Adversarial Attacks Against Large Language Models

ICLR 2026poster

To guarantee safe and robust deployment of large language models (LLMs) at scale, it is critical to accurately assess their adversarial robustness. Existing adversarial attacks typically target harmful responses in single-point greedy generations, overlooking the inherently stochastic nature of LLMs…

Cited by 0SourceScholar