← Search

Mihai Dascalu

7 accepted papers

2025

The Strawberry Problem: Emergence of Character-level Understanding in Tokenized Language Models

EMNLP 2025

Despite their remarkable progress across diverse domains, Large Language Models (LLMs) consistently fail at simple character-level tasks, such as counting letters in words, due to a fundamental limitation: tokenization. In this work, we frame this limitation as a problem of low mutual information an

2024

How Hard is this Test Set? NLI Characterization by Exploiting Training Dynamics

EMNLP 2024main

Natural Language Inference (NLI) evaluation is crucial for assessing language understanding models; however, popular datasets suffer from systematic spurious correlations that artificially inflate actual model performance. To address this, we propose a method for the automated creation of a challeng…

2024

Towards Building the LEMI Readability Platform for Children’s Literature in the Romanian Language

COLING 2024main

Readability is a crucial characteristic of texts, greatly influencing comprehension and reading efficacy. Unfortunately, limited research is available for less-resourced languages, especially for young populations where its impact is even higher. This paper introduces a new readability tool for chil…

Cited by 6SourcePDFScholar
2024

“Vorbești Românește?” A Recipe to Train Powerful Romanian LLMs with English Instructions

EMNLP 2024finding

In recent years, Large Language Models (LLMs) have achieved almost human-like performance on various tasks. While some LLMs have been trained on multilingual data, most of the training data is in English; hence, their performance in English greatly exceeds other languages. To our knowledge, we are t…

2023

TA-DA: Topic-Aware Domain Adaptation for Scientific Keyphrase Identification and Classification (Student Abstract)

AAAI 2023technical

Keyphrase identification and classification is a Natural Language Processing and Information Retrieval task that involves extracting relevant groups of words from a given text related to the main topic. In this work, we focus on extracting keyphrases from scientific documents. We introduce TA-DA, a…

Cited by 1SourcePDFScholar
2022

Domain Adaptation in Multilingual and Multi-Domain Monolingual Settings for Complex Word Identification

ACL 2022long

Complex word identification (CWI) is a cornerstone process towards proper text simplification. CWI is highly dependent on context, whereas its difficulty is augmented by the scarcity of available datasets which vary greatly in terms of domains and languages. As such, it becomes increasingly more dif…

Cited by 4SourcePDFScholar