← Search

Simon Clematide

8 accepted papers

2025

Cheap Character Noise for OCR-Robust Multilingual Embeddings

ACL 2025finding

The large amount of text collections digitized by imperfect OCR systems requires semantic search models that perform robustly on noisy input. Such collections are highly heterogeneous, with varying degrees of OCR quality, spelling conventions and other inconsistencies —all phenomena that are underre…

2025

Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examples

EMNLP 2025

The evaluation of cross-lingual semantic search models is often limited to existing datasets from tasks such as information retrieval and semantic textual similarity. We introduce Cross-Lingual Semantic Discrimination (CLSD), a lightweight evaluation task that requires only parallel sentences and a

2025

Improving Occupational ISCO Classification of Multilingual Swiss Job Postings with LLM-Refined Training Data

ACL 2025finding

Classifying occupations in multilingual job postings is challenging due to noisy labels, language variation, and domain-specific terminology. We present a method that refines silver-standard ISCO labels by consolidating them with predictions from pre-fine-tuned models, using large language model (LL…

Cited by 0SourcePDFScholar
2025

Interpretable Text Embeddings and Text Similarity Explanation: A Survey

EMNLP 2025

Text embeddings are a fundamental component in many NLP tasks, including classification, regression, clustering, and semantic search. However, despite their ubiquitous application, challenges persist in interpreting embeddings and explaining similarities between them.In this work, we provide a struc

Cited by 0SourcePDFScholar
2025

MMTEB: Massive Multilingual Text Embedding Benchmark

ICLR 2025poster

Text embeddings are typically evaluated on a narrow set of tasks, limited in terms of languages, domains, and task types. To circumvent this limitation and to provide a more comprehensive evaluation, we introduce the Massive Multilingual Text Embedding Benchmark (MMTEB) -- a large-scale community-dr…

2025

PARAPHRASUS: A Comprehensive Benchmark for Evaluating Paraphrase Detection Models

COLING 2025main

The task of determining whether two texts are paraphrases has long been a challenge in NLP. However, the prevailing notion of paraphrase is often quite simplistic, offering only a limited view of the vast spectrum of paraphrase phenomena. Indeed, we find that evaluating models in a paraphrase datase…

2025

Sentence Smith: Controllable Edits for Evaluating Text Embeddings

EMNLP 2025

Controllable and transparent text generation has been a long-standing goal in NLP. Almost as long-standing is a general idea for addressing this challenge: Parsing text to a symbolic representation, and generating from it. However, earlier approaches were hindered by parsing and generation insuffici

2024

Mapping Work Task Descriptions from German Job Ads on the O*NET Work Activities Ontology

COLING 2024main

This work addresses the challenge of extracting job tasks from German job postings and mapping them to the fine-grained work activities classification in the O*NET labor market ontology. By utilizing ontological data with a Multiple Negatives Ranking loss and integrating a modest volume of labeled j…

Cited by 0SourcePDFScholar