← Search

Katharina Beckh

2 accepted papers

2025

Robustness Evaluation of the German Extractive Question Answering Task

COLING 2025main

To ensure reliable performance of Question Answering (QA) systems, evaluation of robustness is crucial. Common evaluation benchmarks commonly only include performance metrics, such as Exact Match (EM) and the F1 score. However, these benchmarks overlook critical factors for the deployment of QA syst…

2025

The Anatomy of Evidence: An Investigation Into Explainable ICD Coding

ACL 2025finding

Automatic medical coding has the potential to ease documentation and billing processes. For this task, transparency plays an important role for medical coders and regulatory bodies, which can be achieved using explainability methods. However, the evaluation of these approaches has been mostly limite…

Cited by 0SourcePDFScholar