← Search

Simon Ostermann

9 accepted papers

2025

A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource Languages

EMNLP 2025

Large Language Models (LLMs) are increasingly used to generate synthetic textual data for training smaller specialized models. However, a comparison of various generation strategies for low-resource language settings is lacking. While various prompting strategies have been proposed—such as demonstra

Cited by 0SourcePDFScholar
2025

Cross-Refine: Improving Natural Language Explanation Generation by Learning in Tandem

COLING 2025main

Natural language explanations (NLEs) are vital for elucidating the reasoning behind large language model (LLM) decisions. Many techniques have been developed to generate NLEs using LLMs. However, like humans, LLMs might not always produce optimal NLEs on first attempt. Inspired by human learning pro…

2025

FitCF: A Framework for Automatic Feature Importance-guided Counterfactual Example Generation

ACL 2025finding

Counterfactual examples are widely used in natural language processing (NLP) as valuable data to improve models, and in explainable artificial intelligence (XAI) to understand model behavior. The automated generation of counterfactual examples remains a challenging task even for large language model…

2025

Large Language Models for Multilingual Previously Fact-Checked Claim Detection

EMNLP 2025

In our era of widespread false information, human fact-checkers often face the challenge of duplicating efforts when verifying claims that may have already been addressed in other countries or languages. As false information transcends linguistic boundaries, the ability to automatically detect previ

2024

CoXQL: A Dataset for Parsing Explanation Requests in Conversational XAI Systems

EMNLP 2024finding

Conversational explainable artificial intelligence (ConvXAI) systems based on large language models (LLMs) have garnered significant interest from the research community in natural language processing (NLP) and human-computer interaction (HCI). Such systems can provide answers to user questions abou…

2024

Common European Language Data Space

COLING 2024main

The Common European Language Data Space (LDS) is an integral part of the EU data strategy, which aims at developing a single market for data. Its decentralised technical infrastructure and governance scheme are currently being developed by the LDS project, which also has dedicated tasks for proof-of…

Cited by 2SourcePDFScholar
2024

MMAR: Multilingual and Multimodal Anaphora Resolution in Instructional Videos

EMNLP 2024finding

Multilingual anaphora resolution identifies referring expressions and implicit arguments in texts and links to antecedents that cover several languages. In the most challenging setting, cross-lingual anaphora resolution, training data, and test data are in different languages. As knowledge needs to…

2023

Find-2-Find: Multitask Learning for Anaphora Resolution and Object Localization

EMNLP 2023long main

In multimodal understanding tasks, visual and linguistic ambiguities can arise. Visual ambiguity can occur when visual objects require a model to ground a referring expression in a video without strong supervision, while linguistic ambiguity can occur from changes in entities in action flows. As an…

Cited by 0SourceScholar