← Search

Paula Buttery

9 accepted papers

2025

Rubrik’s Cube: Testing a New Rubric for Evaluating Explanations on the CUBE dataset

ACL 2025long

The performance and usability of Large-Language Models (LLMs) are driving their use in explanation generation tasks. However, despite their widespread adoption, LLM explanations have been found to be unreliable, making it difficult for users to distinguish good from bad explanations. To address this…

2024

Distractor Generation Using Generative and Discriminative Capabilities of Transformer-based Models

COLING 2024main

Multiple Choice Questions (MCQs) are very common in both high-stakes and low-stakes examinations, and their effectiveness in assessing students relies on the quality and diversity of distractors, which are the incorrect answer options provided alongside the correct answer. Motivated by the progress…

Cited by 1SourcePDFScholar
2024

Logging Keystrokes in Writing by English Learners

COLING 2024main

Essay writing is a skill commonly taught and practised in schools. The ability to write a fluent and persuasive essay is often a major component of formal assessment. In natural language processing and education technology we may work with essays in their final form, for example to carry out automat…

Cited by 5SourcePDFScholar
2024

Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic Smoothing

EMNLP 2024main

Language models strongly rely on frequency information because they maximize the likelihood of tokens during pre-training. As a consequence, language models tend to not generalize well to tokens that are seldom seen during training. Moreover, maximum likelihood training has been discovered to give r…

Cited by 1SourcePDFScholar
2024

Prompting open-source and commercial language models for grammatical error correction of English learner text

ACL 2024findings

Thanks to recent advances in generative AI, we are able to prompt large language models (LLMs) to produce texts which are fluent and grammatical. In addition, it has been shown that we can elicit attempts at grammatical error correction (GEC) from LLMs when prompted with ungrammatical input sentence…

Cited by 20SourcePDFScholar
2024

Tending Towards Stability: Convergence Challenges in Small Language Models

EMNLP 2024finding

Increasing the number of parameters in language models is a common strategy to enhance their performance. However, smaller language models remain valuable due to their lower operational costs. Despite their advantages, smaller models frequently underperform compared to their larger counterparts, eve…

2024

Using LLMs to simulate students’ responses to exam questions

EMNLP 2024finding

Previous research leveraged Large Language Models (LLMs) in numerous ways in the educational domain. Here, we show that they can be used to answer exam questions simulating students of different skill levels and share a prompt, engineered for GPT-3.5, that enables the simulation of varying student s…

Cited by 1SourcePDFScholar
2022

Constructing Open Cloze Tests Using Generation and Discrimination Capabilities of Transformers

ACL 2022findings

This paper presents the first multi-objective transformer model for generating open cloze tests that exploits generation and discrimination capabilities to improve performance. Our model is further enhanced by tweaking its loss function and applying a post-processing re-ranking algorithm that improv…

2020

Grammatical error detection in transcriptions of spoken English

COLING 2020main

We describe the collection of transcription corrections and grammatical error annotations for the CrowdED Corpus of spoken English monologues on business topics. The corpus recordings were crowdsourced from native speakers of English and learners of English with German as their first language. The n…