← Search

Jey Han Lau

33 accepted papers

2025

An Interpretable and Crosslingual Method for Evaluating Second-Language Dialogues

NAACL 2025long

We analyse the cross-lingual transferability of a dialogue evaluation framework that assesses the relationships between micro-level linguistic features (e.g. backchannels) and macro-level interactivity labels (e.g. topic management), originally designed for English-as-a-second-language dialogues. To…

2025

Beyond Perception: Evaluating Abstract Visual Reasoning through Multi-Stage Task

ACL 2025finding

Current Multimodal Large Language Models (MLLMs) excel in general visual reasoning but remain underexplored in Abstract Visual Reasoning (AVR), which demands higher-order reasoning to identify abstract rules beyond simple perception. Existing AVR benchmarks focus on single-step reasoning, emphasizin…

2025

Beyond Seen Data: Improving KBQA Generalization Through Schema-Guided Logical Form Generation

EMNLP 2025

Knowledge base question answering (KBQA) aims to answer user questions in natural language using rich human knowledge stored in large KBs. As current KBQA methods struggle with unseen knowledge base elements and their novel compositions at test time, we introduce SG-KBQA — a novel model that injects

2025

Can LLMs Simulate L2-English Dialogue? An Information-Theoretic Analysis of L1-Dependent Biases

ACL 2025long

This study evaluates Large Language Models’ (LLMs) ability to simulate non-native-like English use observed in human second language (L2) learners interfered with by their native first language (L1). In dialogue-based interviews, we prompt LLMs to mimic L2 English learners with specific L1s (e.g., J…

2025

Decomposed Opinion Summarization with Verified Aspect-Aware Modules

ACL 2025finding

Opinion summarization plays a key role in deriving meaningful insights from large-scale online reviews. To make the process more explainable and grounded, we propose a domain-agnostic modular approach guided by review aspects (e.g., cleanliness for hotel reviews) which separates the tasks of aspect…

2025

Evaluating Evidence Attribution in Generated Fact Checking Explanations

NAACL 2025long

Automated fact-checking systems often struggle with trustworthiness, as their generated explanations can include hallucinations. In this work, we explore evidence attribution for fact-checking explanation generation. We introduce a novel evaluation protocol, citation masking and recovery, to assess…

2025

Factual Dialogue Summarization via Learning from Large Language Models

COLING 2025main

Factual consistency is an important quality in dialogue summarization. Large language model (LLM)-based automatic text summarization models generate more factually consistent summaries compared to those by smaller pretrained language models, but they face deployment challenges in real-world applicat…

2025

Interaction Matters: An Evaluation Framework for Interactive Dialogue Assessment on English Second Language Conversations

COLING 2025main

We present an evaluation framework for interactive dialogue assessment in the context of English as a Second Language (ESL) speakers. Our framework collects dialogue-level interactivity labels (e.g., topic management; 4 labels in total) and micro-level span features (e.g., backchannels; 17 features…

2025

Moderation Matters: Measuring Conversational Moderation Impact in English as a Second Language Group Discussion

ACL 2025finding

English as a Second Language (ESL) speakers often struggle to engage in group discussions due to language barriers. While moderators can facilitate participation, few studies assess conversational engagement and evaluate moderation effectiveness. To address this gap, we develop a dataset comprising…

2025

WET: Overcoming Paraphrasing Vulnerabilities in Embeddings-as-a-Service with Linear Transformation Watermarks

ACL 2025long

Embeddings-as-a-Service (EaaS) is a service offered by large language model (LLM) developers to supply embeddings generated by LLMs. Previous research suggests that EaaS is prone to imitation attacks—attacks that clone the underlying EaaS model by training another model on the queried embeddings. As…

2025

WHoW: A Cross-domain Approach for Analysing Conversation Moderation

NAACL 2025long

We propose WHoW, an evaluation framework for analyzing the facilitation strategies of moderators across different domains/scenarios by examining their motives (Why), dialogue acts (How) and target speaker (Who). Using this framework, we annotated 5,657 moderation sentences with human judges and 15,4…

2024

KALE: An Artwork Image Captioning System Augmented with Heterogeneous Graph

IJCAI 2024poster

Exploring the narratives conveyed by fine-art paintings is a challenge in image captioning, where the goal is to generate descriptions that not only precisely represent the visual content but also offer a in-depth interpretation of the artwork's meaning. The task is particularly complex for artwork…

2023

Annotating and Detecting Fine-grained Factual Errors for Dialogue Summarization

ACL 2023long

A series of datasets and models have been proposed for summaries generated for well-formatted documents such as news articles. Dialogue summaries, however, have been under explored. In this paper, we present the first dataset with fine-grained factual error annotations named DIASUMFACT. We define fi…

2023

Compressed Heterogeneous Graph for Abstractive Multi-Document Summarization

AAAI 2023technical

Multi-document summarization (MDS) aims to generate a summary for a number of related documents. We propose HGSum — an MDS model that extends an encoder-decoder architecture to incorporate a heterogeneous graph to represent different semantic units (e.g., words and sentences) of the documents. This…

2023

DeltaScore: Fine-Grained Story Evaluation with Perturbations

EMNLP 2023long findings

Numerous evaluation metrics have been developed for natural language generation tasks, but their effectiveness in evaluating stories is limited as they are not specifically tailored to assess intricate aspects of storytelling, such as fluency and interestingness. In this paper, we introduce DeltaSco…

Cited by 0SourcecodeScholar
2023

Summarizing Multiple Documents with Conversational Structure for Meta-Review Generation

EMNLP 2023long findings

We present PeerSum, a novel dataset for generating meta-reviews of scientific papers. The meta-reviews can be interpreted as abstractive summaries of reviews, multi-turn discussions and the paper abstract. These source documents have a rich inter-document relationship with an explicit hierarchical c…

Cited by 0SourcecodeScholar
2023

Unsupervised Lexical Simplification with Context Augmentation

EMNLP 2023short findings

We propose a new unsupervised lexical simplification method that uses only monolingual data and pre-trained language models. Given a target word and its context, our method generates substitutes based on the target context and also additional contexts sampled from monolingual data. We conduct experi…

Cited by 0SourcecodeScholar
2023

Unsupervised Paraphrasing of Multiword Expressions

ACL 2023findings

We propose an unsupervised approach to paraphrasing multiword expressions (MWEs) in context. Our model employs only monolingual corpus data and pre-trained language models (without fine-tuning), and does not make use of any external resources such as dictionaries. We evaluate our method on the SemEv…

2022

An Interpretable Neuro-Symbolic Reasoning Framework for Task-Oriented Dialogue Generation

ACL 2022long

We study the interpretability issue of task-oriented dialogue systems in this paper. Previously, most neural-based task-oriented dialogue systems employ an implicit reasoning strategy that makes the model predictions uninterpretable to humans. To obtain a transparent reasoning process, we introduce…

2022

DUCK: Rumour Detection on Social Media by Modelling User and Comment Propagation Networks

NAACL 2022long

Social media rumours, a form of misinformation, can mislead the public and cause significant economic and social disruption. Motivated by the observation that the user network — which captures who engage with a story — and the comment network — which captures how they react to it — provide complemen…

2022

LipKey: A Large-Scale News Dataset for Absent Keyphrases Generation and Abstractive Summarization

COLING 2022main

Summaries, keyphrases, and titles are different ways of concisely capturing the content of a document. While most previous work has released the datasets of keyphrases and summarization separately, in this work, we introduce LipKey, the largest news corpus with human-written abstractive summaries, a…

Cited by 9SourcePDFScholar
2022

M3: Multi-level dataset for Multi-document summarisation of Medical studies

EMNLP 2022finding

We present M3 (Multi-level dataset for Multi-document summarisation of Medical studies), a benchmark dataset for evaluating the quality of summarisation systems in the biomedical domain. The dataset contains sets of multiple input documents and target summaries of three levels of complexity: documen…

2022

One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia

ACL 2022long

NLP research is impeded by a lack of resources and awareness of the challenges presented by underrepresented languages and dialects. Focusing on the languages spoken in Indonesia, the second most linguistically diverse and the fourth most populous nation of the world, we provide an overview of the c…

2022

Robust Task-Oriented Dialogue Generation with Contrastive Pre-training and Adversarial Filtering

EMNLP 2022finding

Data artifacts incentivize machine learning models to learn non-transferable generalizations by taking advantage of shortcuts in the data, andthere is growing evidence that data artifacts play a role for the strong results that deep learning models achieve in recent natural language processing bench…

2022

The patient is more dead than alive: exploring the current state of the multi-document summarisation of the biomedical literature

ACL 2022long

Although multi-document summarisation (MDS) of the biomedical literature is a highly valuable task that has recently attracted substantial interest, evaluation of the quality of biomedical summaries lacks consistency and transparency. In this paper, we examine the summaries generated by two current…

Cited by 24SourcePDFScholar
2022

Unsupervised Lexical Substitution with Decontextualised Embeddings

COLING 2022main

We propose a new unsupervised method for lexical substitution using pre-trained language models. Compared to previous approaches that use the generative capability of language models to predict substitutes, our method retrieves substitutes based on the similarity of contextualised and decontextualis…

2021

Automatic Classification of Neutralization Techniques in the Narrative of Climate Change Scepticism

NAACL 2021long

Neutralisation techniques, e.g. denial of responsibility and denial of victim, are used in the narrative of climate change scepticism to justify lack of action or to promote an alternative view. We first draw on social science to introduce the problem to the community of nlp, present the granularity…

Cited by 11SourcePDFScholar
2021

Grey-box Adversarial Attack And Defence For Sentiment Classification

NAACL 2021long

We introduce a grey-box adversarial attack and defence framework for sentiment classification. We address the issues of differentiability, label preservation and input reconstruction for adversarial attack and defence in one unified framework. Our results show that once trained, the attacking model…

2021

IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization

EMNLP 2021main

We present IndoBERTweet, the first large-scale pretrained model for Indonesian Twitter that is trained by extending a monolingually-trained Indonesian BERT model with additive domain-specific vocabulary. We focus in particular on efficient model adaptation under vocabulary mismatch, and benchmark di…

2021

UniMF: A Unified Framework to Incorporate Multimodal Knowledge Bases intoEnd-to-End Task-Oriented Dialogue Systems

IJCAI 2021poster

Knowledge bases (KBs) are usually essential for building practical dialogue systems. Recently we have seen rapidly growing interest in integrating knowledge bases into dialogue systems. However, existing approaches mostly deal with knowledge bases of a single modality, typically textual information.…

2020

IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP

COLING 2020main

Although the Indonesian language is spoken by almost 200 million people and the 10th most spoken language in the world, it is under-represented in NLP research. Previous work on Indonesian has been hampered by a lack of annotated datasets, a sparsity of language resources, and a lack of resource sta…