← Search

Andreas Vlachos

36 accepted papers

2025

A Bayesian Optimization Approach to Machine Translation Reranking

NAACL 2025long

Reranking, or scoring a list of prediction candidates from a machine translation system with an external scoring model and returning the highest-scoring candidate, remains a simple and effective method for improving prediction quality. However, reranking with high quality scoring models can add subs…

2025

AVerImaTeC: A Dataset for Automatic Verification of Image-Text Claims with Evidence from the Web

NeurIPS 2025poster

Textual claims are often accompanied by images to enhance their credibility and spread on social media, but this also raises concerns about the spread of misinformation. Existing datasets for automated verification of image-text claims remain limited, as they often consist of synthetic claims and…

Cited by 0SourceScholar
2025

Causal Estimation of Tokenisation Bias

ACL 2025long

Modern language models are typically trained over subword sequences, but ultimately define probabilities over character-strings. Ideally, the choice of the tokeniser—which maps character-strings to subwords—should not affect the probability assigned to the underlying character-string; in practice, i…

2025

Improving Zero-shot Sentence Decontextualisation with Content Selection and Planning

EMNLP 2025

Extracting individual sentences from a document as evidence or reasoning steps is commonly done in many NLP tasks. However, extracted sentences often lack context necessary to make them understood, e.g., coreference and background information. To this end, we propose a content selection and planning

2025

Introducing FOReCAst: The Future Outcome Reasoning and Confidence Assessment Benchmark

NeurIPS 2025poster

Forecasting is an important task in many domains. However, existing forecasting benchmarks lack comprehensive confidence assessment, focusing on limited question types, and often consist of artificial questions that do not reflect real-world needs. To address these gaps, we introduce FOReCAst (Futur…

Cited by 0SourceScholar
2025

Segment-Level Diffusion: A Framework for Controllable Long-Form Generation with Diffusion Language Models

ACL 2025long

Diffusion models have shown promise in text generation, but often struggle with generating long, coherent, and contextually accurate text. Token-level diffusion doesn’t model word-order dependencies explicitly and operates on short, fixed output windows, while passage-level diffusion struggles with…

Cited by 0SourcePDFScholar
2025

Social Good or Scientific Curiosity? Uncovering the Research Framing Behind NLP Artefacts

EMNLP 2025

Clarifying the research framing of NLP artefacts (e.g., models, datasets, etc.) is crucial to aligning research with practical applications when researchers claim that their findings have real-world impact. Recent studies manually analyzed NLP research across domains, showing that few papers explici

Cited by 0SourcePDFScholar
2025

TCP: a Benchmark for Temporal Constraint-Based Planning

EMNLP 2025

Temporal reasoning and planning are essential capabilities for large language models (LLMs), yet most existing benchmarks evaluate them in isolation and under limited forms of complexity. To address this gap, we introduce the Temporal Constraint-based Planning (TCP) benchmark, that jointly assesses

2024

An LLM Feature-based Framework for Dialogue Constructiveness Assessment

EMNLP 2024main

Research on dialogue constructiveness assessment focuses on (i) analysing conversational factors that influence individuals to take specific actions, win debates, change their perspectives or broaden their open-mindedness and (ii) predicting constructiveness outcomes following dialogues for such use…

2024

AnchorAL: Computationally Efficient Active Learning for Large and Imbalanced Datasets

NAACL 2024long

Active learning for imbalanced classification tasks is challenging as the minority classes naturally occur rarely. Gathering a large pool of unlabelled data is thus essential to capture minority instances. Standard pool-based active learning is computationally expensive on large pools and often reac…

2024

Automated Focused Feedback Generation for Scientific Writing Assistance

ACL 2024findings

Scientific writing is a challenging task, particularly for novice researchers who often rely on feedback from experienced peers. Recent work has primarily focused on improving surface form and style rather than manuscript content. In this paper, we propose a novel task: automated focused feedback ge…

2024

Causal Estimation of Memorisation Profiles

ACL 2024long

Understanding memorisation in language models has practical and societal implications, e.g., studying models’ training dynamics or preventing copyright infringements.Prior work defines memorisation as the causal effect of training with an instance on the model’s ability to predict that instance. Thi…

2024

Do We Need Language-Specific Fact-Checking Models? The Case of Chinese

EMNLP 2024main

This paper investigates the potential benefits of language-specific fact-checking models, focusing on the case of Chinese using CHEF dataset. To better reflect real-world fact-checking, we first develop a novel Chinese document-level evidence retriever, achieving state-of-the-art performance. We the…

2024

Document-level Claim Extraction and Decontextualisation for Fact-Checking

ACL 2024long

Selecting which claims to check is a time-consuming task for human fact-checkers, especially from documents consisting of multiple sentences and containing multiple claims. However, existing claim extraction approaches focus more on identifying and extracting claims from individual sentences, e.g.,…

2024

Zero-Shot Fact Verification via Natural Logic and Large Language Models

EMNLP 2024finding

The recent development of fact verification systems with natural logic has enhanced their explainability by aligning claims with evidence through set-theoretic operators, providing faithful justifications. Despite these advancements, such systems often rely on a large amount of training data annotat…

2023

Multimodal Automated Fact-Checking: A Survey

EMNLP 2023long findings

Misinformation is often conveyed in multiple modalities, e.g. a miscaptioned image. Multimodal misinformation is perceived as more credible by humans, and spreads faster than its text-only counterparts. While an increasing body of research investigates automated fact-checking (AFC), previous survey…

Cited by 0SourcecodeScholar
2023

QA-NatVer: Question Answering for Natural Logic-based Fact Verification

EMNLP 2023long main

Fact verification systems assess a claim's veracity based on evidence. An important consideration in designing them is faithfulness, i.e. generating explanations that accurately reflect the reasoning of the model. Recent works have focused on natural logic, which operates directly on natural languag…

Cited by 0SourcecodeScholar
2023

The Intended Uses of Automated Fact-Checking Artefacts: Why, How and Who

EMNLP 2023long findings

Automated fact-checking is often presented as an epistemic tool that fact-checkers, social media consumers, and other stakeholders can use to fight misinformation. Nevertheless, few papers thoroughly discuss \textit{how}. We document this by analysing 100 highly-cited papers, and annotating epistemi…

Cited by 0SourcecodeScholar
2022

How to disagree well: Investigating the dispute tactics used on Wikipedia

EMNLP 2022main

Disagreements are frequently studied from the perspective of either detecting toxicity or analysing argument structure. We propose a framework of dispute tactics which unifies these two perspectives, as well as other dialogue acts which play a role in resolving disputes, such as asking questions and…

2022

Improving Scheduled Sampling with Elastic Weight Consolidation for Neural Machine Translation

EMNLP 2022finding

Despite strong performance in many sequence-to-sequence tasks, autoregressive models trained with maximum likelihood estimation suffer from exposure bias, i.e. the discrepancy between the ground-truth prefixes used during training and the model-generated prefixes used at inference time. Scheduled sa…

2022

Natural Logic-guided Autoregressive Multi-hop Document Retrieval for Fact Verification

EMNLP 2022main

A key component of fact verification is the evidence retrieval, often from multiple documents. Recent approaches use dense representations and condition the retrieval of each document on the previously retrieved ones. The latter step is performed over all the documents in the collection, requiring s…

2022

Opening up Minds with Argumentative Dialogues

EMNLP 2022finding

Recent research on argumentative dialogues has focused on persuading people to take some action, changing their stance on the topic of discussion, or winning debates. In this work, we focus on argumentative dialogues that aim to open up (rather than change) people’s minds to help them become more un…

Cited by 11SourcePDFScholar
2021

FEVEROUS: Fact Extraction and VERification Over Unstructured and Structured information

NeurIPS 2021poster

Fact verification has attracted a lot of attention in the machine learning and natural language processing communities, as it is one of the key methods for detecting misinformation. Existing large-scale benchmarks for this task have focused mostly on textual sources, i.e. unstructured information, a…

Cited by 270SourcecodeScholar
2021

Leveraging Type Descriptions for Zero-shot Named Entity Recognition and Classification

ACL 2021long

A common issue in real-world applications of named entity recognition and classification (NERC) is the absence of annotated data for the target entity classes during training. Zero-shot learning approaches address this issue by learning models from classes with training data that can predict classes…

Cited by 37SourcePDFScholar