Enhancing Discourse Parsing for Local Structures from Social Media with LLM-Generated Data
Martial Pastor, Nelleke Oostdijk, Patricia Martin-Rodilla, Javier Parapar
Abstract
We explore the use of discourse parsers for extracting a particular discourse structure in a real-world social media scenario. Specifically, we focus on enhancing parser performance through the integration of synthetic data generated by large language models (LLMs). We conduct experiments using a newly developed dataset of 1,170 local RST discourse structures, including 900 synthetic and 270 gold examples, covering three social media platforms: online news comments sections, a discussion forum (Reddit), and a social media messaging platform (Twitter). Our primary goal is to assess the impact of LLM-generated synthetic training data on parser performance in a raw text setting without pre-identified discourse units. While both top-down and bottom-up RST architectures greatly benefit from synthetic data, challenges remain in classifying evaluative discourse structures.
BibTeX
@inproceedings{pastor-etal-2025-enhancing,
title = "Enhancing Discourse Parsing for Local Structures from Social Media with {LLM}-Generated Data",
author = "Pastor, Martial and
Oostdijk, Nelleke and
Martin-Rodilla, Patricia and
Parapar, Javier",
editor = "Rambow, Owen and
Wanner, Leo and
Apidianaki, Marianna and
Al-Khalifa, Hend and
Eugenio, Barbara Di and
Schockaert, Steven",
booktitle = "Proceedings of the 31st International Conference on Computational Linguistics",
month = jan,
year = "2025",
address = "Abu Dhabi, UAE",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.coling-main.584/",
pages = "8739--8748"
}