← Search

Malcolm Slaney

6 accepted papers

2023

Disentangling Speech from Surroundings with Neural Embeddings

ICASSP 2023accepted

We present a method to separate speech signals from noisy environments in the embedding space of a neural audio codec. We introduce a new training procedure that allows our model to produce structured encodings of audio waveforms given by embedding vectors, where one part of the embedding vector rep…

Cited by 19SourceScholar
2022

Multi-Channel Speech Denoising for Machine Ears

ICASSP 2022accepted

This work describes a speech denoising system for machine ears that aims to improve speech intelligibility and the overall listening experience in noisy environments. We recorded approximately 100 hours of audio data with reverberation and moderate environmental noise using a pair of microphone arra…

Cited by 0SourceScholar
2018

Using audio-visual information to understand speaker activity: Tracking active speakers on and off screen

ICASSP 2018accepted

We present a system that associates faces with voices in a video by fusing information from the audio and visual signals. The thesis underlying our work is that an extreme simple approach to generating (weak) speech clusters can be combined with strong visual signals to effectively associate faces a…

Cited by 10SourceScholar
2017

CNN architectures for large-scale audio classification

ICASSP 2017accepted

Convolutional Neural Networks (CNNs) have proven very effective in image classification and show promise for audio. We use various CNN architectures to classify the soundtracks of a dataset of 70M training videos (5.24 million hours) with 30,871 video-level labels. We examine fully connected Deep Ne…

Cited by 3037SourceScholar
2015

Probabilistic features for connecting eye gaze to spoken language understanding

ICASSP 2015accepted

Many users obtain content from a screen and want to make requests of a system based on items that they have seen. Eye-gaze information is a valuable signal in speech recognition and spoken-language understanding (SLU) because it provides context for a user's next utterance-what the user says next is…

Cited by 0SourceScholar