← Search

Juan Pablo Zuluaga Gomez

3 accepted papers

2024

Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper

EMNLP 2024finding

The training of automatic speech recognition (ASR) with little to no supervised data remains an open question. In this work, we demonstrate that streaming Transformer-Transducer (TT) models can be trained from scratch in consumer and accessible GPUs in their entirety with pseudo-labeled (PL) speech…

2024

TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR

EMNLP 2024main

In traditional conversational intelligence from speech, a cascaded pipeline is used, involving tasks such as voice activity detection, diarization, transcription, and subsequent processing with different NLP models for tasks like semantic endpointing and named entity recognition (NER). Our paper int…

2023

End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation

EMNLP 2023long main

Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversations by multiple speakers. In this paper, we tackle single-channel multi-speaker conversational ST with an end-to-end an…

Cited by 0SourcecodeScholar