← Search

Thomas Pellegrini

6 accepted papers

2025

Hierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet Tagging

ICASSP 2025accepted

AudioSet is one of the most used and largest datasets in audio tagging, containing about 2 million audio samples that are manually labeled with 527 event categories organized into an ontology. However, the annotations contain inconsistencies, particularly where categories that should be labeled as p…

Cited by 0SourceScholar
2023

Dilated convolution with learnable spacings

ICLR 2023poster

Recent works indicate that convolutional neural networks (CNN) need large receptive fields (RF) to compete with visual transformers and their attention mechanism. In CNNs, RFs can simply be enlarged by increasing the convolution kernel sizes. Yet the number of trainable parameters, which scales quad…

2021

Comparison of Deep Co-Training and Mean-Teacher Approaches for Semi-Supervised Audio Tagging

ICASSP 2021accepted

Recently, a number of semi-supervised learning (SSL) methods, in the framework of deep learning (DL), were shown to achieve state-of-the-art results on image datasets, while using a (very) limited amount of labeled data. To our knowledge, these approaches adapted and applied to audio data are still…

Cited by 0SourceScholar
2021

Fast Threshold Optimization for Multi-Label Audio Tagging Using Surrogate Gradient Learning

ICASSP 2021accepted

Multi-label audio tagging consists of assigning sets of tags to audio recordings. At inference time, thresholds are applied on the confidence scores outputted by a probabilistic classifier, in order to decide which classes are detected active. In this work, we consider having at disposal a trained c…

Cited by 0SourceScholar
2021

Weakly supervised discourse segmentation for multiparty oral conversations

EMNLP 2021main

Discourse segmentation, the first step of discourse analysis, has been shown to improve results for text summarization, translation and other NLP tasks. While segmentation models for written text tend to perform well, they are not directly applicable to spontaneous, oral conversation, which has ling…

2018

Group emotion recognition strategies for entertainment robots

IROS 2018poster

In this paper, a system to determine the emotion of a group of people via facial expression analysis is proposed for the Waseda Entertainment Robots. General models and standard methods for emotion definition and recognition are briefly described, as well as strategies for computing the group global…

Cited by 36SourceScholar