← Search

Marie Tahon

7 accepted papers

2026

SPARSE AUTOENCODERS MAKE AUDIO FOUNDATION MODELS MORE EXPLAINABLE

ICASSP 2026poster

Audio pretrained models are widely employed to solve various tasks in speech processing, sound event detection, or music information retrieval. However, the representations learned by these models are unclear, and their analysis mainly restricts to linear probing of the hidden representations. In th…

Cited by 0SourcePDFScholar
2024

ALLIES: A Speech Corpus for Segmentation, Speaker Diarization, Speech Recognition and Speaker Change Detection

COLING 2024main

This paper presents ALLIES, a meta corpus which gathers and extends existing French corpora collected from radio and TV shows. The corpus contains 1048 audio files for about 500 hours of speech. Agglomeration of data is always a difficult issue, as the guidelines used to collect, annotate and transc…

2024

An Explainable Proxy Model for Multilabel Audio Segmentation

ICASSP 2024accepted

Audio signal segmentation is a key task for automatic audio indexing. It consists of detecting the boundaries of class-homogeneous segments in the signal. In many applications, explainable AI is a vital process for transparency of decision-making with machine learning. In this paper, we propose an e…

Cited by 0SourceScholar
2024

Annotation of Transition-Relevance Places and Interruptions for the Description of Turn-Taking in Conversations in French Media Content

COLING 2024main

Few speech resources describe interruption phenomena, especially for TV and media content. The description of these phenomena may vary across authors: it thus leaves room for improved annotation protocols. We present an annotation of Transition-Relevance Places (TRP) and Floor-Taking event types on…

Cited by 1SourcePDFScholar
2024

Automatic Speech Interruption Detection: Analysis, Corpus, and System

COLING 2024main

Interruption detection is a new yet challenging task in the field of speech processing. This article presents a comprehensive study on automatic speech interruption detection, from the definition of this task, the assembly of a specialized corpus, and the development of an initial baseline system. W…

Cited by 2SourcePDFScholar
2024

Unsupervised multiple domain translation through controlled Disentanglement in variational autoencoder

ICASSP 2024accepted

Unsupervised Multiple Domain Translation is the task of transforming data from one domain to other domains without having paired data to train the systems. Typically, methods based on Generative Adversarial Networks (GANs) are used to address this task. However, our proposal exclusively relies on a…

Cited by 0SourceScholar
2021

Speaker Embeddings for Diarization of Broadcast Data In The Allies Challenge

ICASSP 2021accepted

Diarization consists in the segmentation of speech signals and the clustering of homogeneous speaker segments. State-of-the-art systems typically operate upon speaker embeddings, such as i-vectors or neural x-vectors, extracted from mel cepstral coefficients (MFCCs) or spectrograms. The recent SincN…

Cited by 0SourceScholar