← Search

Dmitriy Genzel

3 accepted papers

2022

Torchaudio: Building Blocks for Audio and Speech Processing

ICASSP 2022accepted

This document describes version 0.10 of TorchAudio: building blocks for machine learning applications in the audio and speech processing domain. The objective of TorchAudio is to accelerate the development and deployment of machine learning applications for researchers and engineers by providing off…

Cited by 0SourceScholar
2021

A General Multi-Task Learning Framework to Leverage Text Data for Speech to Text Tasks

ICASSP 2021accepted

Attention-based sequence-to-sequence modeling provides a powerful and elegant solution for applications that need to map one sequence to a different sequence. Its success heavily relies on the availability of large amounts of training data. This presents a challenge for speech applications where lab…

Cited by 0SourceScholar
2021

Improving Speech Translation by Understanding and Learning from the Auxiliary Text Translation Task

ACL 2021long

Pretraining and multitask learning are widely used to improve the speech translation performance. In this study, we are interested in training a speech translation model along with an auxiliary text translation task. We conduct a detailed analysis to understand the impact of the auxiliary task on th…