← Search

Robert Litschko

11 accepted papers

2025

Cross-Dialect Information Retrieval: Information Access in Low-Resource and High-Variance Languages

COLING 2025main

A large amount of local and culture-specific knowledge (e.g., people, traditions, food) can only be found in documents written in dialects. While there has been extensive research conducted on cross-lingual information retrieval (CLIR), the field of cross-dialect retrieval (CDIR) has received limite…

2025

Evaluating Large Language Models for Cross-Lingual Retrieval

EMNLP 2025

Multi-stage information retrieval (IR) has become a widely-adopted paradigm in search. While Large Language Models (LLMs) have been extensively evaluated as second-stage reranking models for monolingual IR, a systematic large-scale comparison is still lacking for cross-lingual IR (CLIR). Moreover, w

2025

Make Every Letter Count: Building Dialect Variation Dictionaries from Monolingual Corpora

EMNLP 2025

Dialects exhibit a substantial degree of variation due to the lack of a standard orthography. At the same time, the ability of Large Language Models (LLMs) to process dialects remains largely understudied. To address this gap, we use Bavarian as a case study and investigate the lexical dialect under

2025

Reason to Rote: Rethinking Memorization in Reasoning

EMNLP 2025

Large language models readily memorize arbitrary training instances, such as label noise, yet they perform strikingly well on reasoning tasks. In this work, we investigate how language models memorize label noise, and why such memorization in many cases does not heavily affect generalizable reasonin

Cited by 0SourcePDFScholar
2024

To Know or Not To Know? Analyzing Self-Consistency of Large Language Models under Ambiguity

EMNLP 2024finding

One of the major aspects contributing to the striking performance of large language models (LLMs) is the vast amount of factual knowledge accumulated during pre-training. Yet, many LLMs suffer from self-inconsistency, which raises doubts about their trustworthiness and reliability. This paper focuse…

2024

“Seeing the Big through the Small”: Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?

EMNLP 2024finding

Human label variation (HLV) is a valuable source of information that arises when multiple human annotators provide different labels for valid reasons. In Natural Language Inference (NLI) earlier approaches to capturing HLV involve either collecting annotations from many crowd workers to represent hu…

2023

Boosting Zero-shot Cross-lingual Retrieval by Training on Artificially Code-Switched Data

ACL 2023findings

Transferring information retrieval (IR) models from a high-resource language (typically English) to other languages in a zero-shot fashion has become a widely adopted approach. In this work, we show that the effectiveness of zero-shot rankers diminishes when queries and documents are present in diff…

2023

Establishing Trustworthiness: Rethinking Tasks and Model Evaluation

EMNLP 2023short main

Language understanding is a multi-faceted cognitive capability, which the Natural Language Processing (NLP) community has striven to model computationally for decades. Traditionally, facets of linguistic intelligence have been compartmentalized into tasks with specialized model architectures and cor…

Cited by 0SourceScholar
2023

Vicinal Risk Minimization for Few-Shot Cross-lingual Transfer in Abusive Language Detection

EMNLP 2023long main

Cross-lingual transfer learning from high-resource to medium and low-resource languages has shown encouraging results. However, the scarcity of resources in target languages remains a challenge. In this work, we resort to data augmentation and continual pre-training for domain adaptation to improve…

Cited by 0SourceScholar
2022

Parameter-Efficient Neural Reranking for Cross-Lingual and Multilingual Retrieval

COLING 2022main

State-of-the-art neural (re)rankers are notoriously data-hungry which – given the lack of large-scale training data in languages other than English – makes them rarely used in multilingual and cross-lingual retrieval settings. Current approaches therefore commonly transfer rankers trained on English…

2020

Towards Instance-Level Parser Selection for Cross-Lingual Transfer of Dependency Parsers

COLING 2020main

Current methods of cross-lingual parser transfer focus on predicting the best parser for a low-resource target language globally, that is, “at treebank level”. In this work, we propose and argue for a novel cross-lingual transfer paradigm: instance-level parser selection (ILPS), and present a proof-…

Cited by 4SourcePDFScholar