← Search

Yashesh Gaur

14 accepted papers

2025

Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens

ICASSP 2025accepted

Cascaded speech-to-speech translation systems often suffer from the error accumulation problem and high latency, which is a result of cascaded modules whose inference delays accumulate. In this paper, we propose a transducer-based speech translation model that outputs discrete speech tokens in a low…

Cited by 0SourceScholar
2025

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition

ICASSP 2025accepted

We propose the joint speech translation and recognition (JSTAR) model that leverages the fast-slow cascaded encoder architecture for simultaneous end-to-end automatic speech recognition (ASR) and speech translation (ST). The model is transducer-based and uses a multi-objective training strategy that…

Cited by 0SourceScholar
2024

Leveraging Timestamp Information for Serialized Joint Streaming Recognition and Translation

ICASSP 2024accepted

The growing need for instant spoken language transcription and translation is driven by increased global communication and cross-lingual interactions. This has made offering translations in multiple languages essential for user applications. Traditional approaches to automatic speech recognition (AS…

Cited by 0SourceScholar
2022

Continuous Streaming Multi-Talker ASR with Dual-Path Transducers

ICASSP 2022accepted

Streaming recognition of multi-talker conversations has so far been evaluated only for 2-speaker single-turn sessions. In this paper, we investigate it for multi-turn meetings containing multiple speakers using the Streaming Unmixing and Recognition Transducer (SURT) model, and show that naively ext…

Cited by 0SourceScholar
2022

Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers Using End-to-End Speaker-Attributed ASR

ICASSP 2022accepted

This paper presents Transcribe-to-Diarize, a new approach for neural speaker diarization that uses an end-to-end (E2E) speaker-attributed automatic speech recognition (SA-ASR). The E2E SA-ASR is a joint model that was recently proposed for speaker counting, multi-talker speech recognition, and speak…

Cited by 0SourceScholar
2021

Ensemble Combination between Different Time Segmentations

ICASSP 2021accepted

Hypothesis-level combination between multiple models can often yield gains in speech recognition. However, all models in the ensemble are usually restricted to use the same audio segmentation times. This paper proposes to generalise hypothesis-level combination, allowing the use of different audio s…

Cited by 0SourceScholar
2021

Hypothesis Stitcher for End-to-End Speaker-Attributed ASR on Long-Form Multi-Talker Recordings

ICASSP 2021accepted

An end-to-end (E2E) speaker-attributed automatic speech recognition (SA-ASR) model was proposed recently to jointly perform speaker counting, speech recognition and speaker identification. The model achieved a low speaker-attributed word error rate (SA-WER) for monaural overlapped speech comprising…

Cited by 0SourceScholar
2021

Internal Language Model Training for Domain-Adaptive End-To-End Speech Recognition

ICASSP 2021accepted

The efficacy of external language model (LM) integration with existing end-to-end (E2E) automatic speech recognition (ASR) systems can be improved significantly using the internal language model estimation (ILME) method [1]. In this method, the internal LM score is subtracted from the score obtained…

Cited by 0SourceScholar
2021

Minimum Bayes Risk Training for End-to-End Speaker-Attributed ASR

ICASSP 2021accepted

Recently, an end-to-end speaker-attributed automatic speech recognition (E2E SA-ASR) model was proposed as a joint model of speaker counting, speech recognition and speaker identification for monaural overlapped speech. In the previous study, the model parameters were trained based on the speaker-at…

Cited by 0SourceScholar
2020

Minimum Latency Training Strategies for Streaming Sequence-to-Sequence ASR

ICASSP 2020accepted

Recently, a few novel streaming attention-based sequence-to-sequence (S2S) models have been proposed to perform online speech recognition with linear-time decoding complexity. However, in these models, the decisions to generate tokens are delayed compared to the actual acoustic boundaries since thei…

Cited by 0SourceScholar
2018

Robust Speech Recognition Using Generative Adversarial Networks

ICASSP 2018accepted

This paper describes a general, scalable, end-to-end framework that uses the generative adversarial network (GAN) objective to enable robust speech recognition. Encoders trained with the proposed approach enjoy improved invariance by learning to map noisy audio to the same embedding space as that of…

Cited by 0SourceScholar