← Search

Radu Horaud

19 accepted papers

2023

Audio-Visual Speaker Diarization in the Framework of Multi-User Human-Robot Interaction

ICASSP 2023accepted

The speaker diarization task answers the question "who is speaking at a given time?". It represents valuable information for scene analysis in a domain such as robotics. In this paper, we introduce a temporal audio-visual fusion model for multiusers speaker diarization, with low computing requiremen…

Cited by 0SourceScholar
2022

The Impact of Removing Head Movements on Audio-Visual Speech Enhancement

ICASSP 2022accepted

This paper investigates the impact of head movements on audio-visual speech enhancement (AVSE). Although being a common conversational feature, head movements have been ignored by past and recent studies: they challenge today’s learning-based methods as they often degrade the performance of models t…

Cited by 0SourceScholar
2021

Fullsubnet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement

ICASSP 2021accepted

This paper proposes a full-band and sub-band fusion model, named as FullSubNet, for single-channel real-time speech enhancement. Full-band and sub-band refer to the models that input full-band and sub-band noisy spectral feature, output full-band and sub-band speech target, respectively. The sub-ban…

Cited by 0SourceScholar
2020

A Recurrent Variational Autoencoder for Speech Enhancement

ICASSP 2020accepted

This paper presents a generative approach to speech enhancement based on a recurrent variational autoencoder (RVAE). The deep generative speech model is trained using clean speech signals only, and it is combined with a nonnegative matrix factorization noise model for speech enhancement. We propose…

Cited by 0SourceScholar
2020

How to Train Your Deep Multi-Object Tracker

CVPR 2020poster

The recent trend in vision-based multi-object tracking (MOT) is heading towards leveraging the representational power of deep learning to jointly learn to detect and track objects. However, existing methods train only certain sub-modules using loss functions that often do not correlate with establis…

Cited by 274PDFcodeScholar
2019

Semi-supervised Multichannel Speech Enhancement with Variational Autoencoders and Non-negative Matrix Factorization

ICASSP 2019accepted

In this paper we address speaker-independent multichannel speech enhancement in unknown noisy environments. Our work is based on a well-established multichannel local Gaussian modeling framework. We propose to use a neural network for modeling the speech spectro-temporal content. The parameters of t…

Cited by 0SourceScholar
2019

Speech Enhancement with Variational Autoencoders and Alpha-stable Distributions

ICASSP 2019accepted

This paper focuses on single-channel semi-supervised speech enhancement. We learn a speaker-independent deep generative speech model using the framework of variational autoencoders. The noise model remains unsupervised because we do not assume prior knowledge of the noisy recording environment. In t…

Cited by 0SourceScholar
2018

Accounting for Room Acoustics in Audio-Visual Multi-Speaker Tracking

ICASSP 2018accepted

Multiple-speaker tracking is a crucial task for many applications. In real-world scenarios, exploiting the complementarity between auditory and visual data enables to track people outside the visual field of view. However, practical methods must be robust to changes in acoustic conditions, e.g. reve…

Cited by 0SourceScholar
2018

Deep Reinforcement Learning for Audio-Visual Gaze Control

IROS 2018poster

We address the problem of audio-visual gaze control in the specific context of human-robot interaction, namely how controlled robot motions are combined with visual and acoustic observations in order to direct the robot head towards targets of interest. The paper has the following contributions: (i)…

Cited by 18SourceScholar
2018

DeepGUM: Learning Deep Robust Regression with a Gaussian-Uniform Mixture Model

ECCV 2018poster

In this paper we address the problem of how to robustly train a ConvNet for regression, or deep robust regression. Traditionally, deep regression employ the L2 loss function, known to be sensitive to outliers, i.e. samples that either lie at an abnormal distance away from the majority of the trainin…

Cited by 37SourcePDFScholar
2017

An EM algorithm for joint source separation and diarisation of multichannel convolutive speech mixtures

ICASSP 2017accepted

We present a probabilistic model for joint source separation and diarisation of multichannel convolutive speech mixtures. We build upon the framework of local Gaussian model (LGM) with non-negative matrix factorization (NMF). The diarisation is introduced as a temporal labeling of each source in the…

Cited by 0SourceScholar
2017

Audio source separation based on convolutive transfer function and frequency-domain lasso optimization

ICASSP 2017accepted

This paper addresses the problem of under-determined convolutive audio source separation in a semi-oracle configuration where the mixing filters are assumed to be known. We propose a separation procedure based on the convolutive transfer function (CTF), which is a more appropriate model for strongly…

Cited by 0SourceScholar
2017

Deep Mixture of Linear Inverse Regressions Applied to Head-Pose Estimation

CVPR 2017poster

Convolutional Neural Networks (ConvNets) have become the state-of-the-art for many classification and regression problems in computer vision. When it comes to regression, approaches such as measuring the Euclidean distance of target and predictions are often employed as output layer. In this paper,…

Cited by 69PDFScholar
2017

Tracking a varying number of people with a visually-controlled robotic head

IROS 2017poster

Multi-person tracking with a robotic platform is one of the cornerstones of human-robot interaction. Challenges arise from occlusions, appearance changes and a time-varying number of people. Furthermore, the final system is constrained by the hardware platform: low computational capacity and limited…

Cited by 25SourceScholar
2016

An inverse-gamma source variance prior with factorized parameterization for audio source separation

ICASSP 2016accepted

In this paper we present a new statistical model for the power spectral density (PSD) of an audio signal and its application to multichannel audio source separation (MASS). The source signal is modeled with the local Gaussian model (LGM) and we propose to model its variance with an inverse-Gamma dis…

Cited by 0SourceScholar
2016

Non-stationary noise power spectral density estimation based on regional statistics

ICASSP 2016accepted

Estimating the noise power spectral density (PSD) is essential for single channel speech enhancement algorithms. In this paper, we propose a noise PSD estimation approach based on regional statistics. The proposed regional statistics consist of four features representing the statistics of the past a…

Cited by 0SourceScholar
2016

Reverberant sound localization with a robot head based on direct-path relative transfer function

IROS 2016poster

This paper addresses the problem of sound-source localization (SSL) with a robot head, which remains a challenge in real-world environments. In particular we are interested in locating speech sources, as they are of high interest for human-robot interaction. The microphone-pair response correspondin…

Cited by 44SourceScholar
2015

Estimation of relative transfer function in the presence of stationary noise based on segmental power spectral density matrix subtraction

ICASSP 2015accepted

This paper addresses the problem of relative transfer function (RTF) estimation in the presence of stationary noise. We propose an RTF identification method based on segmental power spectral density (PSD) matrix subtraction. First multiple channel microphone signals are divided into segments corresp…

Cited by 0SourceScholar