← Search

John Pavlopoulos

14 accepted papers

2025

Dialect Normalization using Large Language Models and Morphological Rules

ACL 2025finding

Natural language understanding systems struggle with low-resource languages, including many dialects of high-resource ones. Dialect-to-standard normalization attempts to tackle this issue by transforming dialectal text so that it can be used by standard-language tools downstream. In this study, we t…

2025

Evaluation and Facilitation of Online Discussions in the LLM Era: A Survey

EMNLP 2025

We present a survey of methods for assessing and enhancing the quality of online discussions, focusing on the potential of Large Language Models (LLMs). While online discourses aim, at least in theory, to foster mutual understanding, they often devolve into harmful exchanges, such as hate speech, th

Cited by 0SourcePDFScholar
2025

FoodSafeSum: Enabling Natural Language Processing Applications for Food Safety Document Summarization and Analysis

EMNLP 2025

Food safety demands timely detection, regulation, and public communication, yet the lack of structured datasets hinders Natural Language Processing (NLP) research. We present and release a new dataset of human-written and Large Language Model (LLM)-generated summaries of food safety documents, plus

Cited by 0SourcePDFScholar
2025

GR-NLP-TOOLKIT: An Open-Source NLP Toolkit for Modern Greek

COLING 2025system demonstrations

We present GR-NLP-TOOLKIT, an open-source natural language processing (NLP) toolkit developed specifically for modern Greek. The toolkit provides state-of-the-art performance in five core NLP tasks, namely part-of-speech tagging, morphological tagging, dependency parsing, named entity recognition, a…

2025

Learning to Align: Addressing Character Frequency Distribution Shifts in Handwritten Text Recognition

EMNLP 2025

Handwritten text recognition aims to convert visual input into machine-readable text, and it remains challenging due to the evolving and context-dependent nature of handwriting. Character sets change over time, and character frequency distributions shift across historical periods or regions, often c

2024

CICLe: Conformal In-Context Learning for Largescale Multi-Class Food Risk Classification

ACL 2024findings

Contaminated or adulterated food poses a substantial risk to human health. Given sets of labeled web texts for training, Machine Learning and Natural Language Processing can be applied to automatically detect such risks. We publish a dataset of 7,546 short texts describing public food recall announc…

2024

Deciphering Emotional Landscapes in the Iliad: A Novel French-Annotated Dataset for Emotion Recognition

COLING 2024main

One of the most significant pieces of ancient Greek literature, the Iliad, is part of humanity’s collective cultural heritage. This work aims to provide the scientific community with an emotion-labeled dataset for classical literature and Western mythology in particular. To model the emotions of the…

Cited by 1SourcePDFScholar
2024

HoLM: Analyzing the Linguistic Unexpectedness in Homeric Poetry

COLING 2024main

The authorship of the Homeric poems has been a matter of debate for centuries. Computational approaches such as language modeling exist that can aid experts in making crucial headway. We observe, however, that such work has, thus far, only been carried out at the level of lengthier excerpts, but not…

2024

Still All Greeklish to Me: Greeklish to Greek Transliteration

COLING 2024main

Modern Greek is normally written in the Greek alphabet. In informal online messages, however, Greek is often written using characters available on Latin-character keyboards, a form known as Greeklish. Originally used to bypass the lack of support for the Greek alphabet in older computers, Greeklish…

2024

Towards a Greek Proverb Atlas: Computational Spatial Exploration and Attribution of Greek Proverbs

EMNLP 2024main

Proverbs carry wisdom transferred orally from generation to generation. Based on the place they were recorded, this study introduces a publicly-available and machine-actionable dataset of more than one hundred thousand Greek proverb variants. By quantifying the spatial distribution of proverbs, we s…

2023

Dating Greek Papyri with Text Regression

ACL 2023long

Dating Greek papyri accurately is crucial not only to edit their texts but also to understand numerous other aspects of ancient writing, document and book production and circulation, as well as various other aspects of administration, everyday life and intellectual history of antiquity. Although a s…

2023

Detecting Erroneously Recognized Handwritten Byzantine Text

EMNLP 2023long findings

Handwritten text recognition (HTR) yields textual output that comprises errors, which are considerably more compared to that of recognised printed (OCRed) text. Post-correcting methods can eliminate such errors but may also introduce errors. In this study, we investigate the issues arising from this…

Cited by 0SourceScholar
2022

A Greek Parliament Proceedings Dataset for Computational Linguistics and Political Analysis

NeurIPS 2022accept

Large, diachronic datasets of political discourse are hard to come across, especially for resource-lean languages such as Greek. In this paper, we introduce a curated dataset of the Greek Parliament Proceedings that extends chronologically from 1989 up to 2020. It consists of more than 1 million spe…

2022

From the Detection of Toxic Spans in Online Discussions to the Analysis of Toxic-to-Civil Transfer

ACL 2022long

We study the task of toxic spans detection, which concerns the detection of the spans that make a text toxic, when detecting such spans is possible. We introduce a dataset for this task, ToxicSpans, which we release publicly. By experimenting with several methods, we show that sequence labeling mode…