← Search

Bjoern Eskofier

5 accepted papers

2025

Stratify or Die: Rethinking Data Splits in Image Segmentation

NeurIPS 2025poster

Random splitting of datasets in image segmentation often leads to unrepresentative test sets, resulting in biased evaluations and poor model generalization. While stratified sampling has proven effective for addressing label distribution imbalance in classification tasks, extending these ideas to se…

Cited by 0SourceScholar
2024

On the Scalability of Certified Adversarial Robustness with Generated Data

NeurIPS 2024poster

Certified defenses against adversarial attacks offer formal guarantees on the robustness of a model, making them more reliable than empirical methods such as adversarial training, whose effectiveness is often later reduced by unseen attacks. Still, the limited certified robustness that is currently…

Cited by 0SourcePDFScholar
2023

$p$-value Adjustment for Monotonous, Unbiased, and Fast Clustering Comparison

NeurIPS 2023poster

Popular metrics for clustering comparison, like the Adjusted Rand Index and the Adjusted Mutual Information, are type II biased. The Standardized Mutual Information removes this bias but suffers from counterintuitive non-monotonicity and poor computational efficiency. We introduce the $p$-value adju…

2022

Improving Robustness against Real-World and Worst-Case Distribution Shifts through Decision Region Quantification

ICML 2022spotlight

The reliability of neural networks is essential for their use in safety-critical applications. Existing approaches generally aim at improving the robustness of neural networks to either real-world distribution shifts (e.g., common corruptions and perturbations, spatial transformations, and natural a…

Cited by 20SourcePDFScholar
2021

Identifying untrustworthy predictions in neural networks by geometric gradient analysis

UAI 2021poster

The susceptibility of deep neural networks to untrustworthy predictions, including out-of-distribution (OOD) data and adversarial examples, still prevent their widespread use in safety-critical applications. Most existing methods either require a retraining of a given model to achieve robust identif…

Cited by 16SourcePDFScholar