← Search

Djamé Seddah

7 accepted papers

2026

Disentangling meaning from language in LLM-based machine translation

ICML 2026poster

Mechanistic Interpretability (MI) seeks to explain how neural networks implement their capabilities, but the scale of Large Language Models (LLMs) has limited prior MI work in Machine Translation (MT) to word-level analyses. We study sentence-level MT from a mechanistic perspective by analyzing atte…

Cited by 0SourceScholar
2025

Beyond Dataset Creation: Critical View of Annotation Variation and Bias Probing of a Dataset for Online Radical Content Detection

COLING 2025main

The proliferation of radical content on online platforms poses significant risks, including inciting violence and spreading extremist ideologies. Despite ongoing research, existing datasets and models often fail to address the complexities of multilingual and diverse data. To bridge this gap, we int…

Cited by 0SourcePDFScholar
2024

From Text to Source: Results in Detecting Large Language Model-Generated Content

COLING 2024main

The widespread use of Large Language Models (LLMs), celebrated for their ability to generate human-like text, has raised concerns about misinformation and ethical implications. Addressing these concerns necessitates the development of robust methods to detect and attribute text generated by LLMs. Th…

Cited by 11SourcePDFScholar
2022

Exploiting Inductive Bias in Transformers for Unsupervised Disentanglement of Syntax and Semantics with VAEs

NAACL 2022long

We propose a generative model for text generation, which exhibits disentangled latent representations of syntax and semantics. Contrary to previous work, this model does not need syntactic information such as constituency parses, or semantic information such as paraphrase pairs. Our model relies sol…

2021

Synthetic Data Augmentation for Zero-Shot Cross-Lingual Question Answering

EMNLP 2021main

Coupled with the availability of large scale datasets, deep learning architectures have enabled rapid progress on the Question Answering task. However, most of those datasets are in English, and the performances of state-of-the-art multilingual models are significantly lower when evaluated on non-En…

2021

When Being Unseen from mBERT is just the Beginning: Handling New Languages With Multilingual Language Models

NAACL 2021long

Transfer learning based on pretraining language models on a large amount of raw data has become a new norm to reach state-of-the-art performance in NLP. Still, it remains unclear how this approach should be applied for unseen languages that are not covered by any available large-scale multilingual l…

Cited by 153SourcePDFScholar