← Search

Leo Liu

5 accepted papers

2024

Personalization of CTC-Based End-to-End Speech Recognition Using Pronunciation-Driven Subword Tokenization

ICASSP 2024accepted

Recent advances in deep learning and automatic speech recognition have improved the accuracy of end-to-end speech recognition systems, but recognition of personal content such as contact names remains a challenge. In this work, we describe our personalization solution for an end-to-end speech recogn…

Cited by 0SourceScholar
2023

Training Large-Vocabulary Neural Language Models by Private Federated Learning for Resource-Constrained Devices

ICASSP 2023accepted

Federated Learning (FL) is a technique to train models on distributed edge devices with local data samples. Differential Privacy (DP) can be applied with FL to provide a formal privacy guarantee for sensitive data on device. Our goal is to train a large neural network language model (NNLM) on comput…

Cited by 0SourceScholar
2022

Development of a Stingray-inspired High-Frequency Propulsion Platform with Variable Wavelength

IROS 2022poster

Undulatory fin motions in fish-like robots are typically created using intricate arrays of servo motors. Motor arrays offer impressive versatility in terms of kinematics, but their complexity leads to constraints on size, hydrodynamic force production, and power consumption, particularly when studyi…

Cited by 4SourceScholar
2020

SNDCNN: Self-Normalizing Deep CNNs with Scaled Exponential Linear Units for Speech Recognition

ICASSP 2020accepted

Very deep CNNs achieve state-of-the-art results in both computer vision and speech recognition, but are difficult to train. The most popular way to train very deep CNNs is to use shortcut connections (SC) together with batch normalization (BN). Inspired by Self-Normalizing Neural Networks, we propos…

Cited by 41SourceScholar
2019

Voice Trigger Detection from Lvcsr Hypothesis Lattices Using Bidirectional Lattice Recurrent Neural Networks

ICASSP 2019accepted

We propose a method to reduce false voice triggers of a speech-enabled personal assistant by post-processing the hypothesis lattice of a server-side large-vocabulary continuous speech recognizer (LVCSR) via a neural network. We first discuss how an estimate of the posterior probability of the trigge…

Cited by 0SourceScholar