← Search

Verena Blaschke

7 accepted papers

2025

Analyzing the Effect of Linguistic Similarity on Cross-Lingual Transfer: Tasks and Experimental Setups Matter

ACL 2025finding

Cross-lingual transfer is a popular approach to increase the amount of training data for NLP tasks in a low-resource context. However, the best strategy to decide which cross-lingual data to include is unclear. Prior research often focuses on a small set of languages from a few language families and…

2025

Cross-Dialect Information Retrieval: Information Access in Low-Resource and High-Variance Languages

COLING 2025main

A large amount of local and culture-specific knowledge (e.g., people, traditions, food) can only be found in documents written in dialects. While there has been extensive research conducted on cross-lingual information retrieval (CLIR), the field of cross-dialect retrieval (CDIR) has received limite…

2025

Evaluating Pixel Language Models on Non-Standardized Languages

COLING 2025main

We explore the potential of pixel-based models for transfer learning from standard languages to dialects. These models convert text into images that are divided into patches, enabling a continuous vocabulary representation that proves especially useful for out-of-vocabulary words common in dialectal…

Cited by 1SourcePDFScholar
2025

Make Every Letter Count: Building Dialect Variation Dictionaries from Monolingual Corpora

EMNLP 2025

Dialects exhibit a substantial degree of variation due to the lack of a standard orthography. At the same time, the ability of Large Language Models (LLMs) to process dialects remains largely understudied. To address this gap, we use Bavarian as a case study and investigate the lexical dialect under

2024

MaiBaam: A Multi-Dialectal Bavarian Universal Dependency Treebank

COLING 2024main

Despite the success of the Universal Dependencies (UD) project exemplified by its impressive language breadth, there is still a lack in ‘within-language breadth’: most treebanks focus on standard languages. Even for German, the language with the most annotations in UD, so far no treebank exists for…

2024

Sebastian, Basti, Wastl?! Recognizing Named Entities in Bavarian Dialectal Data

COLING 2024main

Named Entity Recognition (NER) is a fundamental task to extract key information from texts, but annotated resources are scarce for dialects. This paper introduces the first dialectal NER dataset for German, BarNER, with 161K tokens annotated on Bavarian Wikipedia articles (bar-wiki) and tweets (bar-…

2024

What Do Dialect Speakers Want? A Survey of Attitudes Towards Language Technology for German Dialects

ACL 2024short

Natural language processing (NLP) has largely focused on modelling standardized languages. More recently, attention has increasingly shifted to local, non-standardized languages and dialects. However, the relevant speaker populations’ needs and wishes with respect to NLP tools are largely unknown. I…

Cited by 12SourcePDFScholar