← Search

Max W. Y. Lam

11 accepted papers

2024

SongCreator: Lyrics-based Universal Song Generation

NeurIPS 2024poster

Music is an integral part of human culture, embodying human intelligence and creativity, of which songs compose an essential part. While various aspects of song generation have been explored by previous works, such as singing voice, vocal composition and instrumental arrangement, etc., generating so…

2023

Efficient Neural Music Generation

NeurIPS 2023poster

Recent progress in music generation has been remarkably advanced by the state-of-the-art MusicLM, which comprises a hierarchy of three LMs, respectively, for semantic, coarse acoustic, and fine acoustic modelings. Yet, sampling with the MusicLM requires processing through these LMs one by one to obt…

2022

BDDM: Bilateral Denoising Diffusion Models for Fast and High-Quality Speech Synthesis

ICLR 2022poster

Diffusion probabilistic models (DPMs) and their extensions have emerged as competitive generative models yet confront challenges of efficient sampling. We propose a new bilateral denoising diffusion model (BDDM) that parameterizes both the forward and reverse processes with a schedule network and a…

2022

FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis

IJCAI 2022poster

Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hindered their applications to speech synthesis. This paper proposes FastDiff, a fast conditional diffusion model for high-qu…

2021

Sandglasset: A Light Multi-Granularity Self-Attentive Network for Time-Domain Speech Separation

ICASSP 2021accepted

One of the leading single-channel speech separation (SS) models is based on a TasNet with a dual-path segmentation technique, where the size of each segment remains unchanged throughout all layers. In contrast, our key finding is that multi-granularity features are essential for enhancing contextual…

Cited by 0SourceScholar
2021

Tune-In: Training Under Negative Environments with Interference for Attention Networks Simulating Cocktail Party Effect

AAAI 2021technical

We study the cocktail party problem and propose a novel attention network called Tune-In, abbreviated for training under negative environments with interference. It firstly learns two separate spaces of speaker-knowledge and speech-stimuli based on a shared feature space, where a new block structure…

2020

Mixup-breakdown: A Consistency Training Method for Improving Generalization of Speech Separation Models

ICASSP 2020accepted

Deep-learning based speech separation models confront poor generalization problem that even the state-of-the-art models could abruptly fail when evaluating them in mismatch conditions. To address this problem, we propose an easy-to-implement yet effective consistency based semi-supervised learning (…

Cited by 0SourceScholar
2019

Bayesian and Gaussian Process Neural Networks for Large Vocabulary Continuous Speech Recognition

ICASSP 2019accepted

The hidden activation functions inside deep neural networks (DNNs) play a vital role in learning high level discriminative features and controlling the information flows to track longer history. However, the fixed model parameters used in standard DNNs can lead to over-fitting and poor generalizatio…

Cited by 0SourceScholar
2019

Gaussian Process Lstm Recurrent Neural Network Language Models for Speech Recognition

ICASSP 2019accepted

Recurrent neural network language models (RNNLMs) have shown superior performance across a range of speech recognition tasks. At the heart of all RNNLMs, the activation functions play a vital role to control the information flows and tracking longer history contexts that are useful for predicting th…

Cited by 0SourceScholar
2019

Recurrent Neural Network Language Model Training Using Natural Gradient

ICASSP 2019accepted

Recurrent neural network language models (RNNLMs) have become an increasing popular choice for state-of-the-art speech recognition systems. RNNLMs are normally trained by minimizing the cross entropy (CE) using the stochastic gradient descent (SGD) algorithm. However, the SGD method doesn't consider…

Cited by 0SourceScholar