← Search

Andros Tjandra

14 accepted papers

2024

Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of a Multilingual ASR Model

ICASSP 2024accepted

Neural network pruning offers an effective method for compressing a multilingual automatic speech recognition (ASR) model with minimal performance loss. However, it entails several rounds of pruning and re-training needed to be run for each language. In this work, we propose the use of an adaptive m…

Cited by 0SourceScholar
2024

Generative Pre-training for Speech with Flow Matching

ICLR 2024poster

Generative models have gained more and more attention in recent years for their remarkable success in tasks that required estimating and sampling data distribution to generate high-fidelity synthetic data. In speech, text-to-speech synthesis and neural vocoder are good examples where generative mode…

Cited by 34SourcePDFScholar
2024

MusicFlow: Cascaded Flow Matching for Text Guided Music Generation

ICML 2024poster

We introduce MusicFlow, a cascaded text-to-music generation model based on flow matching. Based on self-supervised representations to bridge between text descriptions and music audios, we construct two flow matching networks to model the conditional distribution of semantic and acoustic features. Ad…

Cited by 9SourcePDFScholar
2023

Learning ASR Pathways: A Sparse Multilingual ASR Model

ICASSP 2023accepted

Neural network pruning compresses automatic speech recognition (ASR) models effectively. However, in multilingual ASR, language-agnostic pruning may lead to severe performance drops on some languages because language-agnostic pruning masks may not fit all languages and discard important language-spe…

Cited by 0SourceScholar
2023

Massively Multilingual ASR on 70 Languages: Tokenization, Architecture, and Generalization Capabilities

ICASSP 2023accepted

End-to-end multilingual ASR has become more appealing because of several reasons such as simplifying the training and deployment process and positive performance transfer from high-resource to low-resource languages. However, scaling up the number of languages, total hours, and number of unique toke…

Cited by 0SourceScholar
2023

Voice-Preserving Zero-Shot Multiple Accent Conversion

ICASSP 2023accepted

Most people who have tried to learn a foreign language would have experienced difficulties understanding or speaking with a native speaker’s accent. For native speakers, understanding or speaking a new accent is likewise a difficult task. An accent conversion system that changes a speaker’s accent b…

Cited by 26SourceScholar
2022

Conformer-Based Self-Supervised Learning For Non-Speech Audio Tasks

ICASSP 2022accepted

Representation learning from unlabeled data has been of major interest in artificial intelligence research. While self-supervised speech representation learning has been popular in the speech research community, very few works have comprehensively analyzed audio representation learning for non-speec…

Cited by 0SourceScholar
2022

Improved Language Identification Through Cross-Lingual Self-Supervised Learning

ICASSP 2022accepted

Language identification greatly impacts the success of downstream tasks such as automatic speech recognition. Recently, self-supervised speech representations learned by wav2vec 2.0 have been shown to be very effective for a range of speech tasks. We extend previous self-supervised work on language…

Cited by 0SourceScholar
2020

DEJA-VU: Double Feature Presentation and Iterated Loss in Deep Transformer Networks

ICASSP 2020accepted

Deep acoustic models typically receive features in the first layer of the network, and process increasingly abstract representations in the subsequent layers. Here, we propose to feed the input features at multiple depths in the acoustic model. As our motivation is to allow acoustic models to re-exa…

Cited by 0SourceScholar
2020

Transformer-Based Acoustic Modeling for Hybrid Speech Recognition

ICASSP 2020accepted

We propose and evaluate transformer-based acoustic models (AMs) for hybrid speech recognition. Several modeling choices are discussed in this work, including various positional embedding methods and an iterated loss to enable training deep transformers. We also present a preliminary study of using l…

Cited by 0SourceScholar
2019

End-to-end Feedback Loss in Speech Chain Framework via Straight-through Estimator

ICASSP 2019accepted

The speech chain mechanism integrates automatic speech recognition (ASR) and text-to-speech synthesis (TTS) modules into a single cycle during training. In our previous work, we applied a speech chain mechanism as a semi-supervised learning. It provides the ability for ASR and TTS to assist each oth…

Cited by 0SourceScholar
2015

Combination of two-dimensional cochleogram and spectrogram features for deep learning-based ASR

ICASSP 2015accepted

This paper explores the use of auditory features based on cochleograms; two dimensional speech features derived from gammatone filters within the convolutional neural network (CNN) framework. Furthermore, we also propose various possibilities to combine cochleogram features with log-mel filter banks…

Cited by 22SourceScholar