← Search

Antonios Anastasopoulos

38 accepted papers

2025

Cross-Lingual Representation Alignment Through Contrastive Image-Caption Tuning

ACL 2025short

Multilingual alignment of sentence representations has mostly required bitexts to bridge the gap between languages. We investigate whether visual information can bridge this gap instead. Image caption datasets are very easy to create without requiring multilingual expertise, so this offers a more ef…

2025

Dialect Normalization using Large Language Models and Morphological Rules

ACL 2025finding

Natural language understanding systems struggle with low-resource languages, including many dialects of high-resource ones. Dialect-to-standard normalization attempts to tackle this issue by transforming dialectal text so that it can be used by standard-language tools downstream. In this study, we t…

2025

Dialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varieties

EMNLP 2025

There has been little systematic study on how dialectal differences affect toxicity detection by modern LLMs. Furthermore, although using LLMs as evaluators (“LLM-as-a-judge”) is a growing research area, their sensitivity to dialectal nuances is still underexplored and requires more focused attentio

2025

Follow the Beaten Path: The Role of Route Patterns on Vision-Language Navigation Agents Generalization Abilities

NAACL 2025long

Vision and language navigation (VLN) is a challenging task towards the creation of embodied agents that requires spatial and temporal reasoning over the instructions provided in natural language and aligning them with the visual perception of an environment. Although a number of methods and approach…

2025

Script-Agnosticism and its Impact on Language Identification for Dravidian Languages

NAACL 2025long

Language identification is used as the first step in many data collection and crawling efforts because it allows us to sort online text into language-specific buckets. However, many modern languages, such as Konkani, Kashmiri, Punjabi etc., are synchronically written in several scripts. Moreover, la…

2025

Tracing L1 Interference in English Learner Writing: A Longitudinal Corpus with Error Annotations

EMNLP 2025

Language transfer is an important topic of research in second language acquisition and computational linguistics. The availability of suitable learner corpora is paramount for the study of second language acquisition (SLA) and language transfer. However, curating learner corpora is a challenging end

2025

mHumanEval - A Multilingual Benchmark to Evaluate Large Language Models for Code Generation

NAACL 2025long

Recent advancements in large language models (LLMs) have significantly enhanced code generation from natural language prompts. The HumanEval Benchmark, developed by OpenAI, remains the most widely used code generation benchmark. However, this and other Code LLM benchmarks face critical limitations,…

2024

BiasDora: Exploring Hidden Biased Associations in Vision-Language Models

EMNLP 2024finding

Existing works examining Vision-Language Models (VLMs) for social biases predominantly focus on a limited set of documented bias associations, such as gender-profession or race-crime. This narrow scope often overlooks a vast range of unexamined implicit associations, restricting the identification a…

2024

Birdie: Advancing State Space Language Modeling with Dynamic Mixtures of Training Objectives

EMNLP 2024main

Efficient state space models (SSMs), including linear recurrent neural networks and linear attention variants, have emerged as potential alternative language models to Transformers. While efficient, SSMs struggle with tasks requiring in-context retrieval, such as text copying and associative recall,…

2024

Cross-Modal Multi-Tasking for Speech-to-Text Translation via Hard Parameter Sharing

ICASSP 2024accepted

Recent works in end-to-end speech-to-text translation (ST) have proposed multi-tasking methods with soft parameter sharing which leverage machine translation (MT) data via secondary encoders that map text inputs to an eventual cross-modal representation. In this work, we instead propose a ST/MT mult…

Cited by 0SourceScholar
2024

DIALECTBENCH: An NLP Benchmark for Dialects, Varieties, and Closely-Related Languages

ACL 2024long

Language technologies should be judged on their usefulness in real-world use cases. An often overlooked aspect in natural language processing (NLP) research and evaluation is language variation in the form of non-standard dialects or language varieties (hereafter, varieties). Most NLP benchmarks are…

2024

Dictionary-Aided Translation for Handling Multi-Word Expressions in Low-Resource Languages

ACL 2024findings

Multi-word expressions (MWEs) present unique challenges in natural language processing (NLP), particularly within the context of translation systems, due to their inherent scarcity, non-compositional nature, and other distinct lexical and morphosyntactic characteristics, issues that are exacerbated…

2024

Enhancing End-to-End Conversational Speech Translation Through Target Language Context Utilization

ICASSP 2024accepted

Incorporating longer context has been shown to benefit machine translation, but the inclusion of context in end-to-end speech translation (E2E-ST) remains under-studied. To bridge this gap, we introduce target language context in E2E-ST, enhancing coherence and overcoming memory constraints of exten…

Cited by 0SourceScholar
2024

Extracting Lexical Features from Dialects via Interpretable Dialect Classifiers

NAACL 2024short

Identifying linguistic differences between dialects of a language often requires expert knowledge and meticulous human analysis. This is largely due to the complexity and nuance involved in studying various dialects. We present a novel approach to extract distinguishing lexical features of dialects…

2024

Global Gallery: The Fine Art of Painting Culture Portraits through Multilingual Instruction Tuning

NAACL 2024long

Exploring the intersection of language and culture in Large Language Models (LLMs), this study critically examines their capability to encapsulate cultural nuances across diverse linguistic landscapes. Central to our investigation are three research questions: the efficacy of language-specific instr…

2024

Gloss2Text: Sign Language Gloss translation using LLMs and Semantically Aware Label Smoothing

EMNLP 2024finding

Sign language translation from video to spoken text presents unique challenges owing to the distinct grammar, expression nuances, and high variation of visual appearance across different speakers and contexts. Gloss annotations serve as an intermediary to guide the translation process. In our work,…

2024

Language and Speech Technology for Central Kurdish Varieties

COLING 2024main

Kurdish, an Indo-European language spoken by over 30 million speakers, is considered a dialect continuum and known for its diversity in language varieties. Previous studies addressing language and speech technology for Kurdish handle it in a monolithic way as a macro-language, resulting in dispariti…

2024

The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead?

EMNLP 2024main

Large Language Models (LLMs) have shown capabilities close to human performance in various analytical tasks, leading researchers to use them for time and labor-intensive analyses. However, their capability to handle highly specialized and open-ended tasks in domains like policy studies remains in qu…

2023

BIG-C: a Multimodal Multi-Purpose Dataset for Bemba

ACL 2023long

We present BIG-C (Bemba Image Grounded Conversations), a large multimodal dataset for Bemba. While Bemba is the most populous language of Zambia, it exhibits a dearth of resources which render the development of language technologies or language processing research almost impossible. The dataset is…

2023

Global Voices, Local Biases: Socio-Cultural Prejudices across Languages

EMNLP 2023long main

Human biases are ubiquitous but not uniform: disparities exist across linguistic, cultural, and societal borders. As large amounts of recent literature suggest, language models (LMs) trained on human data can reflect and often amplify the effects of these social biases. However, the vast majority of…

Cited by 0SourcecodeScholar
2023

GlobalBench: A Benchmark for Global Progress in Natural Language Processing

EMNLP 2023long main

Despite the major advances in NLP, significant disparities in NLP system performance across languages still exist. Arguably, these are due to uneven resource allocation and sub-optimal incentives to work on less resourced languages. To track and further incentivize the global development of equitabl…

Cited by 0SourceScholar
2023

LIMIT: Language Identification, Misidentification, and Translation using Hierarchical Models in 350+ Languages

EMNLP 2023long main

Knowing the language of an input text/audio is a necessary first step for using almost every NLP tool such as taggers, parsers, or translation systems. Language identification is a well-studied problem, sometimes even considered solved; in reality, due to lack of data and computational challenges, c…

Cited by 0SourcecodeScholar
2023

Script Normalization for Unconventional Writing of Under-Resourced Languages in Bilingual Communities

ACL 2023long

The wide accessibility of social media has provided linguistically under-represented communities with an extraordinary opportunity to create content in their native languages. This, however, comes with certain challenges in script normalization, particularly where the speakers of a language in a bil…

2023

Teacher Perception of Automatically Extracted Grammar Concepts for L2 Language Learning

EMNLP 2023long findings

One of the challenges in language teaching is how best to organize rules regarding syntax, semantics, or phonology in a meaningful manner. This not only requires content creators to have pedagogical skills, but also have that language's deep understanding. While comprehensive materials to develop…

Cited by 0SourceScholar
2022

Revisiting the Effects of Leakage on Dependency Parsing

ACL 2022findings

Recent work by Søgaard (2020) showed that, treebank size aside, overlap between training and test graphs (termed leakage) explains more of the observed variation in dependency parsing performance than other explanations. In this work we revisit this claim, testing it on more models and languages. We…

2022

Systematic Inequalities in Language Technology Performance across the World’s Languages

ACL 2022long

Natural language processing (NLP) systems have become a central technology in communication, education, medicine, artificial intelligence, and many other domains of research and development. While the performance of NLP methods has grown enormously over the last decade, this progress has been restri…

2021

Evaluating the Morphosyntactic Well-formedness of Generated Texts

EMNLP 2021main

Text generation systems are ubiquitous in natural language processing applications. However, evaluation of these systems remains a challenge, especially in multilingual settings. In this paper, we propose L’AMBRE – a metric to evaluate the morphosyntactic well-formedness of text using its dependency…

2021

Machine Translation into Low-resource Language Varieties

ACL 2021short

State-of-the-art machine translation (MT) systems are typically trained to generate “standard” target language; however, many languages have multiple varieties (regional varieties, dialects, sociolects, non-native varieties) that are different from the standard language. Such varieties are often low…

2021

SD-QA: Spoken Dialectal Question Answering for the Real World

EMNLP 2021finding

Question answering (QA) systems are now available through numerous commercial applications for a wide variety of domains, serving millions of users that interact with them via speech interfaces. However, current benchmarks in QA research do not account for the errors that speech recognition models m…

2021

Towards more equitable question answering systems: How much more data do you need?

ACL 2021short

Question answering (QA) in English has been widely explored, but multilingual datasets are relatively new, with several methods attempting to bridge the gap between high- and low-resourced languages using data augmentation through translation and cross-lingual transfer. In this project we take a ste…

2021

When Being Unseen from mBERT is just the Beginning: Handling New Languages With Multilingual Language Models

NAACL 2021long

Transfer learning based on pretraining language models on a large amount of raw data has become a new norm to reach state-of-the-art performance in NLP. Still, it remains unclear how this approach should be applied for unseen languages that are not covered by any available large-scale multilingual l…

Cited by 153SourcePDFScholar
2021

When is Wall a Pared and when a Muro?: Extracting Rules Governing Lexical Selection

EMNLP 2021main

Learning fine-grained distinctions between vocabulary items is a key challenge in learning a new language. For example, the noun “wall” has different lexical manifestations in Spanish – “pared” refers to an indoor wall while “muro” refers to an outside wall. However, this variety of lexical distinct…

Cited by 3SourcePDFScholar
2020

Automatic Interlinear Glossing for Under-Resourced Languages Leveraging Translations

COLING 2020main

Interlinear Glossed Text (IGT) is a widely used format for encoding linguistic information in language documentation projects and scholarly papers. Manual production of IGT takes time and requires linguistic expertise. We attempt to address this issue by creating automatic glossing models, using mod…

2020

Optimizing Data Usage via Differentiable Rewards

ICML 2020poster

To acquire a new skill, humans learn better and faster if a tutor, based on their current knowledge level, informs them of how much attention they should pay to particular content or practice problems. Similarly, a machine learning model could potentially be trained better with a scorer that “adapts…

2020

Universal Phone Recognition with a Multilingual Allophone System

ICASSP 2020accepted

Multilingual models can improve language processing, particularly for low resource situations, by sharing parameters across languages. Multilingual acoustic models, however, generally ignore the difference between phonemes (sounds that can support lexical contrasts in a particular language) and thei…

Cited by 0SourceScholar