← Search

Gregory Sell

11 accepted papers

2020

A Practical Two-Stage Training Strategy for Multi-Stream End-to-End Speech Recognition

ICASSP 2020accepted

The multi-stream paradigm of audio processing, in which several sources are simultaneously considered, has been an active research area for information fusion. Our previous study offered a promising direction within end-to-end automatic speech recognition, where parallel encoders aim to capture dive…

Cited by 0SourceScholar
2020

Jhu-HLTCOE System for the Voxsrc Speaker Recognition Challenge

ICASSP 2020accepted

The VoxSRC speaker recognition challenge comprises data obtained from YouTube videos of celebrity interviews in a wide range of recording environments. The challenge provides FIXED and OPEN training conditions to allow cross-system comparisons and to characterize the effects of additional amounts of…

Cited by 0SourceScholar
2019

Deriving Spectro-temporal Properties of Hearing from Speech Data

ICASSP 2019accepted

Human hearing and human speech are intrinsically tied together, as the properties of speech almost certainly developed in order to be heard by human ears. As a result of this connection, it has been shown that certain properties of human hearing are mimicked within data-driven systems that are train…

Cited by 0SourceScholar
2019

Joint Acoustic and Class Inference for Weakly Supervised Sound Event Detection

ICASSP 2019accepted

Sound event detection is a challenging task, especially for scenes with multiple simultaneous events. While event classification methods tend to be fairly accurate, event localization presents additional challenges, especially when large amounts of labeled data are not available. Task4 of the 2018 D…

Cited by 0SourceScholar
2019

Speaker Recognition for Multi-speaker Conversations Using X-vectors

ICASSP 2019accepted

Recently, deep neural networks that map utterances to fixed-dimensional embeddings have emerged as the state-of-the-art in speaker recognition. Our prior work introduced x-vectors, an embedding that is very effective for both speaker recognition and diarization. This paper combines our previous work…

Cited by 0SourceScholar
2018

Audio-Visual Person Recognition in Multimedia Data From the Iarpa Janus Program

ICASSP 2018accepted

Currently, datasets that support audio-visual recognition of people in videos are scarce and limited. In this paper, we introduce an expansion of video data from the IARPA Janus program to support this research area. We refer to the expanded set, which adds labels for voice to the already-existing f…

Cited by 0SourceScholar
2018

X-Vectors: Robust DNN Embeddings for Speaker Recognition

ICASSP 2018accepted

In this paper, we use data augmentation to improve performance of deep neural network (DNN) embeddings for speaker recognition. The DNN, which is trained to discriminate between speakers, maps variable-length utterances to fixed-dimensional embeddings that we call x-vectors. Prior studies have found…

Cited by 0SourceScholar
2017

Speaker diarization using deep neural network embeddings

ICASSP 2017accepted

Speaker diarization is an important front-end for many speech technologies in the presence of multiple speakers, but current methods that employ i-vector clustering for short segments of speech are potentially too cumbersome and costly for the front-end role. In this work, we propose an alternative…

Cited by 0SourceScholar
2015

Content-based recommender systems for spoken documents

ICASSP 2015accepted

Content-based recommender systems use preference ratings and features that characterize media to model users' interests or information needs for making future recommendations. While previously developed in the music and text domains, we present an initial exploration of content-based recommendation…

Cited by 16SourceScholar