← Search

Lucia Specia

19 accepted papers

2025

DiffuseDef: Improved Robustness to Adversarial Attacks via Iterative Denoising

ACL 2025long

Pretrained language models have significantly advanced performance across various natural language processing tasks. However, adversarial attacks continue to pose a critical challenge to system built using these models, as they can be exploited with carefully crafted adversarial texts. Inspired by t…

2024

Classifying Social Media Users before and after Depression Diagnosis via Their Language Usage: A Dataset and Study

COLING 2024main

Mental illness can significantly impact individuals’ quality of life. Analysing social media data to uncover potential mental health issues in individuals via their posts is a popular research direction. However, most studies focus on the classification of users suffering from depression versus heal…

2022

Logically Consistent Adversarial Attacks for Soft Theorem Provers

IJCAI 2022poster

Recent efforts within the AI community have yielded impressive results towards “soft theorem proving” over natural language sentences using language models. We propose a novel, generative adversarial framework for probing and improving these models’ reasoning capabilities. Adversarial attacks in thi…

2021

BERTGen: Multi-task Generation through BERT

ACL 2021long

We present BERTGen, a novel, generative, decoder-only model which extends BERT by fusing multimodal and multilingual pre-trained models VL-BERT and M-BERT, respectively. BERTGen is auto-regressively trained for language generation tasks, namely image captioning, machine translation and multimodal ma…

2021

Backtranslation Feedback Improves User Confidence in MT, Not Quality

NAACL 2021long

Translating text into a language unknown to the text’s author, dubbed outbound translation, is a modern need for which the user experience has significant room for improvement, beyond the basic machine translation facility. We demonstrate this by showing three ways in which user confidence in the ou…

2021

Classification-based Quality Estimation: Small and Efficient Models for Real-world Applications

EMNLP 2021main

Sentence-level Quality estimation (QE) of machine translation is traditionally formulated as a regression task, and the performance of QE models is typically measured by Pearson correlation with human labels. Recent QE models have achieved previously-unseen levels of correlation with human judgments…

Cited by 3SourcePDFScholar
2021

deepQuest-py: Large and Distilled Models for Quality Estimation

EMNLP 2021system demonstrations

We introduce deepQuest-py, a framework for training and evaluation of large and light-weight models for Quality Estimation (QE). deepQuest-py provides access to (1) state-of-the-art models based on pre-trained Transformers for sentence-level and word-level QE; (2) light-weight and efficient sentence…

2020

Curious Case of Language Generation Evaluation Metrics: A Cautionary Tale

COLING 2020main

Automatic evaluation of language generation systems is a well-studied problem in Natural Language Processing. While novel metrics are proposed every year, a few popular metrics remain as the de facto metrics to evaluate tasks such as image captioning and machine translation, despite their known limi…

2016

Groupwise learning for ASR k-best list reranking in spoken language translation

ICASSP 2016accepted

Quality estimation models are used to predict the quality of the output from a spoken language translation (SLT) system. When these scores are used to rerank a k-best list, the rank of the scores is more important than their absolute values. This paper proposes groupwise learning to model this rank.…

Cited by 0SourceScholar
2015

Quality estimation for asr k-best list rescoring in spoken language translation

ICASSP 2015accepted

Spoken language translation (SLT) combines automatic speech recognition (ASR) and machine translation (MT). During the decoding stage, the best hypothesis produced by the ASR system may not be the best input candidate to the MT system, but making use of multiple sub-optimal ASR results in SLT has be…

Cited by 0SourceScholar