← Search

Cristina España-Bonet

13 accepted papers

2026

Grounding or Guessing? Visual Signals for Detecting Hallucinations in Sign Language Translation

ICLR 2026poster

Hallucination, where models generate fluent text unsupported by visual evidence, remains a major flaw in vision–language models and is especially critical in sign language translation (SLT). In SLT, meaning depends on precise grounding in video, and gloss-free models are particularly vulnerable beca…

Cited by 0SourceScholar
2026

REVISITING DIRECT SPEECH-TO-TEXT TRANSLATION WITH SPEECH LLMS: BETTER SCALING THAN COT PROMPTING?

ICASSP 2026poster

Recent work on Speech-to-Text Translation (S2TT) has focused on LLM-based models, introducing the increasingly adopted Chain-of-Thought (CoT) prompting, where the model is guided to first transcribe the speech and then translate it. CoT typically outperforms direct prompting primarily because it can…

Cited by 0SourcePDFScholar
2025

Continual Learning in Multilingual Sign Language Translation

NAACL 2025long

The field of sign language translation (SLT) is still in its infancy, as evidenced by the low translation quality, even when using deep learn- ing approaches. Probably because of this, many common approaches in other machine learning fields have not been explored in sign language. Here, we focus on…

Cited by 0SourcePDFScholar
2024

DGS-Fabeln-1: A Multi-Angle Parallel Corpus of Fairy Tales between German Sign Language and German Text

COLING 2024main

We present the acquisition process and the data of DGS-Fabeln-1, a parallel corpus of German text and videos containing German fairy tales interpreted into the German Sign Language (DGS) by a native DGS signer. The corpus contains 573 segments of videos with a total duration of 1 hour and 32 minutes…

Cited by 1SourcePDFScholar
2024

Sign Language Translation with Sentence Embedding Supervision

ACL 2024short

State-of-the-art sign language translation (SLT) systems facilitate the learning process through gloss annotations, either in an end2end manner or by involving an intermediate step. Unfortunately, gloss labelled sign language data is usually not available at scale and, when available, gloss annotati…

2024

When Your Cousin Has the Right Connections: Unsupervised Bilingual Lexicon Induction for Related Data-Imbalanced Languages

COLING 2024main

Most existing approaches for unsupervised bilingual lexicon induction (BLI) depend on good quality static or contextual embeddings requiring large monolingual corpora for both languages. However, unsupervised BLI is most likely to be useful for low-resource languages (LRLs), where large datasets are…

2023

Multilingual Coarse Political Stance Classification of Media. The Editorial Line of a ChatGPT and Bard Newspaper

EMNLP 2023short findings

Neutrality is difficult to achieve and, in politics, subjective. Traditional media typically adopt an editorial line that can be used by their potential readers as an indicator of the media bias. Several platforms currently rate news outlets according to their political bias. The editorial line and…

Cited by 0SourceScholar
2023

Translating away Translationese without Parallel Data

EMNLP 2023long main

Translated texts exhibit systematic linguistic differences compared to original texts in the same language, and these differences are referred to as translationese. Translationese has effects on various cross-lingual natural language processing tasks, potentially leading to biased results. In this p…

Cited by 0SourceScholar
2022

The (Undesired) Attenuation of Human Biases by Multilinguality

EMNLP 2022main

Some human preferences are universal. The odor of vanilla is perceived as pleasant all around the world. We expect neural models trained on human texts to exhibit these kind of preferences, i.e. biases, but we show that this is not always the case. We explore 16 static and contextual embedding model…

2022

Towards Debiasing Translation Artifacts

NAACL 2022long

Cross-lingual natural language processing relies on translation, either by humans or machines, at different levels, from translating training data to translating test sets. However, compared to original texts in the same language, translations possess distinct qualities referred to as translationese…

2021

Comparing Feature-Engineering and Feature-Learning Approaches for Multilingual Translationese Classification

EMNLP 2021main

Traditional hand-crafted linguistically-informed features have often been used for distinguishing between translated and original non-translated texts. By contrast, to date, neural architectures without manual feature engineering have been less explored for this task. In this work, we (i) compare th…

Cited by 21SourcePDFScholar
2020

Understanding Translationese in Multi-view Embedding Spaces

COLING 2020main

Recent studies use a combination of lexical and syntactic features to show that footprints of the source language remain visible in translations, to the extent that it is possible to predict the original source language from the translation. In this paper, we focus on embedding-based semantic spaces…

Cited by 14SourcePDFScholar