← Search

Tannon Kew

4 accepted papers

2025

Robust Native Language Identification through Agentic Decomposition

EMNLP 2025

Large language models (LLMs) often achieve high performance in native language identification (NLI) benchmarks by leveraging superficial contextual clues such as names, locations, and cultural stereotypes, rather than the underlying linguistic patterns indicative of native language (L1) influence. T

2024

Turning English-centric LLMs Into Polyglots: How Much Multilinguality Is Needed?

EMNLP 2024finding

The vast majority of today’s large language models (LLMs) are English-centric, having been pretrained predominantly on English text. Yet, in order to meet user expectations, models need to be able to respond appropriately in multiple languages once deployed in downstream applications. This requires…

2023

BLESS: Benchmarking Large Language Models on Sentence Simplification

EMNLP 2023long main

We present BLESS, a comprehensive performance benchmark of the most recent state-of-the-art Large Language Models (LLMs) on the task of text simplification (TS). We examine how well off-the-shelf LLMs can solve this challenging task, assessing a total of 44 models, differing in size, architecture, p…

Cited by 0SourcecodeScholar
2023

Uncovering Hidden Consequences of Pre-training Objectives in Sequence-to-Sequence Models

ACL 2023findings

Some variants of self-supervised denoising objectives for pre-training encoder-decoder language models have been reported to have a negligible impact on downstream performance. Yet the design of these pre-training objectives leads to behavioural differences that can be uncovered with specific manipu…