← Search

Ehsan Variani

12 accepted papers

2025

Massive Sound Embedding Benchmark (MSEB)

NeurIPS 2025poster

Audio is a critical component of multimodal perception, and any truly intelligent system must demonstrate a wide range of auditory capabilities. These capabilities include transcription, classification, retrieval, reasoning, segmentation, clustering, reranking, and reconstruction. Fundamentally, eac…

Cited by 0SourcecodeScholar
2023

JEIT: Joint End-to-End Model and Internal Language Model Training for Speech Recognition

ICASSP 2023accepted

We propose JEIT, a joint end-to-end (E2E) model and internal language model (ILM) training method to inject large-scale unpaired text into ILM during E2E training which improves rare-word speech recognition. With JEIT, the E2E model computes an E2E loss on audio-transcript pairs while its ILM estima…

Cited by 0SourceScholar
2022

Global Normalization for Streaming Speech Recognition in a Modular Framework

NeurIPS 2022accept

We introduce the Globally Normalized Autoregressive Transducer (GNAT) for addressing the label bias problem in streaming speech recognition. Our solution admits a tractable exact computation of the denominator for the sequence-level normalization. Through theoretical and empirical results, we demons…

Cited by 11SourcePDFScholar
2022

Multilingual Second-Pass Rescoring for Automatic Speech Recognition Systems

ICASSP 2022accepted

Second-pass rescoring is a well known technique to improve the performance of Automatic Speech Recognition (ASR) systems. Neural Oracle Search (NOS), which selects the most likely hypothesis from an N-best hypothesis list by integrating information from multiple sources, such as the input acoustic r…

Cited by 0SourceScholar
2021

Cascaded Encoders for Unifying Streaming and Non-Streaming ASR

ICASSP 2021accepted

End-to-end (E2E) automatic speech recognition (ASR) models, by now, have shown competitive performance on several benchmarks. These models are structured to either operate in streaming or non-streaming mode. This work presents cascaded encoders for building a single E2E ASR model that can operate in…

Cited by 0SourceScholar
2020

Neural Oracle Search on N-BEST Hypotheses

ICASSP 2020accepted

In this paper, we propose a neural search algorithm to select the most likely hypothesis using a sequence of acoustic representations and multiple hypotheses as input. The algorithm provides a sequence level score for each audio-hypothesis pair that is obtained by integrating information from multip…

Cited by 0SourceScholar
2018

Sampled Connectionist Temporal Classification

ICASSP 2018accepted

This article introduces and evaluates Sampled Connectionist Temporal Classification (CTC) which connects the CTC criterion to the Cross Entropy (CE) objective through sampling. Instead of computing the logarithm of the sum of the alignment path likelihoods, at each training step the sampled CTC only…

Cited by 0SourceScholar
2015

A Gaussian Mixture Model layer jointly optimized with discriminative features within a Deep Neural Network architecture

ICASSP 2015accepted

This article proposes and evaluates a Gaussian Mixture Model (GMM) represented as the last layer of a Deep Neural Network (DNN) architecture and jointly optimized with all previous layers using Asynchronous Stochastic Gradient Descent (ASGD). The resulting “Deep GMM” architecture was investigated wi…

Cited by 0SourceScholar