← Search

Elena Simperl

8 accepted papers

2026

TSM-Bench: Detecting LLM-Generated Text in Real-World Wikipedia Editing Practices

ICLR 2026poster

Automatically detecting machine-generated text (MGT) is critical to maintaining the knowledge integrity of user-generated content (UGC) platforms such as Wikipedia. Existing detection benchmarks primarily focus on \textit{generic} text generation tasks (e.g., ``Write an article about machine learni…

Cited by 0SourceScholar
2025

Schema Generation for Large Knowledge Graphs Using Large Language Models

EMNLP 2025

Schemas play a vital role in ensuring data quality and supporting usability in the Semantic Web and natural language processing. Traditionally, their creation demands substantial involvement from knowledge engineers and domain experts. Leveraging the impressive capabilities of large language models

2024

ChartCheck: Explainable Fact-Checking over Real-World Chart Images

ACL 2024findings

Whilst fact verification has attracted substantial interest in the natural language processing community, verifying misinforming statements against data visualizations such as charts has so far been overlooked. Charts are commonly used in the real-world to summarize and com municate key information,…

2024

Croissant: A Metadata Format for ML-Ready Datasets

NeurIPS 2024spotlight

Data is a critical resource for machine learning (ML), yet working with data remains a key friction point. This paper introduces Croissant, a metadata format for datasets that creates a shared representation across ML tools, frameworks, and platforms. Croissant makes datasets more discoverable, por…

2023

Exploring the Numerical Reasoning Capabilities of Language Models: A Comprehensive Analysis on Tabular Data

EMNLP 2023long findings

Numerical data plays a crucial role in various real-world domains like finance, economics, and science. Thus, understanding and reasoning with numbers are essential in these fields. Recent benchmarks have assessed the numerical reasoning abilities of language models, revealing their limitations in l…

Cited by 0SourceScholar
2023

Multimodal Automated Fact-Checking: A Survey

EMNLP 2023long findings

Misinformation is often conveyed in multiple modalities, e.g. a miscaptioned image. Multimodal misinformation is perceived as more credible by humans, and spreads faster than its text-only counterparts. While an increasing body of research investigates automated fact-checking (AFC), previous survey…

Cited by 0SourcecodeScholar
2022

PubHealthTab: A Public Health Table-based Dataset for Evidence-based Fact Checking

NAACL 2022findings

Inspired by human fact checkers, who use different types of evidence (e.g. tables, images, audio) in addition to text, several datasets with tabular evidence data have been released in recent years. Whilst the datasets encourage research on table fact-checking, they rely on information from restrict…

2020

Point at the Triple: Generation of Text Summaries from Knowledge Base Triples (Extended Abstract)

IJCAI 2020poster

We investigate the problem of generating natural language summaries from knowledge base triples. Our approach is based on a pointer-generator network, which, in addition to generating regular words from a fixed target vocabulary, is able to verbalise triples in several ways. We undertake an automati…