← Search

Wassim Bouaziz

7 accepted papers

2026

Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset

ICLR 2026poster

How can large language models (LLMs) serve users with varying preferences that may conflict across cultural, political, or other dimensions? To advance this challenge, this paper establishes four key results. First, we demonstrate, through a large-scale multilingual human study with representative s…

Cited by 0SourcecodeScholar
2026

Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data Poisoning

ICLR 2026poster

The pre-training of large language models (LLMs) relies on massive text datasets sourced from diverse and difficult-to-curate origins. Although membership inference attacks and hidden canaries have been explored to trace data usage, such methods rely on *regurgitation* of training data, which LM pro…

Cited by 0SourceScholar
2025

Data Taggants: Dataset Ownership Verification Via Harmless Targeted Data Poisoning

ICLR 2025poster

Dataset ownership verification, the process of determining if a dataset is used in a model's training data, is necessary for detecting unauthorized data usage and data contamination. Existing approaches, such as backdoor watermarking, rely on inducing a detectable behavior into the trained model on…

Cited by 2SourcePDFScholar
2025

Targeted Data Poisoning for Black-Box Audio Datasets Ownership Verification

ICASSP 2025accepted

Protecting the use of audio datasets is a major concern for data owners, particularly with the recent rise of audio deep learning models. While watermarks can be used to protect the data itself, they do not allow to identify a deep learning model trained on a protected dataset. In this paper, we ada…

Cited by 0SourceScholar
2024

Iteration Head: A Mechanistic Study of Chain-of-Thought

NeurIPS 2024poster

Chain-of-Thought (CoT) reasoning is known to improve Large Language Models both empirically and in terms of theoretical approximation power. However, our understanding of the inner workings and conditions of apparition of CoT capabilities remains limited. This paper helps fill this gap by demonstrat…

2020

Pyannote.Audio: Neural Building Blocks for Speaker Diarization

ICASSP 2020accepted

We introduce pyannote.audio, an open-source toolkit written in Python for speaker diarization. Based on PyTorch machine learning framework, it provides a set of trainable end-to-end neural building blocks that can be combined and jointly optimized to build speaker diarization pipelines. pyannote.aud…

Cited by 476SourceScholar