← Search

Philippe Muller

6 accepted papers

2024

DISRPT: A Multilingual, Multi-domain, Cross-framework Benchmark for Discourse Processing

COLING 2024main

This paper presents DISRPT, a multilingual, multi-domain, and cross-framework benchmark dataset for discourse processing, covering the tasks of discourse unit segmentation, connective identification, and relation classification. DISRPT includes 13 languages, with data from 24 corpora covering about…

2024

In2Core: Leveraging Influence Functions for Coreset Selection in Instruction Finetuning of Large Language Models

EMNLP 2024finding

Despite advancements, fine-tuning Large Language Models (LLMs) remains costly due to the extensive parameter count and substantial data requirements for model generalization. Accessibility to computing resources remains a barrier for the open-source community. To address this challenge, we propose t…

Cited by 1SourcePDFScholar
2024

Zero-shot Learning for Multilingual Discourse Relation Classification

COLING 2024main

Classifying discourse relations is known as a hard task, relying on complex indices. On the other hand, discourse-annotated data is scarce, especially for languages other than English: many corpora, of limited size, exist for several languages but the domain is split between different theoretical fr…

2023

An Integrated Approach for Political Bias Prediction and Explanation Based on Discursive Structure

ACL 2023findings

One crucial aspect of democracy is fair information sharing. While it is hard to prevent biases in news, they should be identified for better transparency. We propose an approach to automatically characterize biases that takes into account structural differences and that is efficient for long texts.…

2023

Leveraging Argumentation for Generating Robust Sample-based Explanations

IJCAI 2023poster

Explaining predictions made by inductive classifiers has become crucial with the rise of complex models acting more and more as black-boxes. Abductive explanations are one of the most popular types of explanations that are provided for the purpose. They highlight feature-values that are suffici…

Cited by 6SourcePDFScholar
2021

Weakly supervised discourse segmentation for multiparty oral conversations

EMNLP 2021main

Discourse segmentation, the first step of discourse analysis, has been shown to improve results for text summarization, translation and other NLP tasks. While segmentation models for written text tend to perform well, they are not directly applicable to spontaneous, oral conversation, which has ling…