← Search

Qiujia Li

11 accepted papers

2024

Efficient Adapter Finetuning for Tail Languages in Streaming Multilingual ASR

ICASSP 2024accepted

The end-to-end ASR model is often desired in the streaming multilingual scenario since it is easier to deploy and can benefit from pre-trained speech models such as powerful foundation models. Meanwhile, the heterogeneous nature and imbalanced data abundance of different languages may cause performa…

Cited by 0SourceScholar
2024

Handling Ambiguity in Emotion: From Out-of-Domain Detection to Distribution Estimation

ACL 2024long

The subjective perception of emotion leads to inconsistent labels from human annotators. Typically, utterances lacking majority-agreed labels are excluded when training an emotion classifier, which cause problems when encountering ambiguous emotional expressions during testing. This paper investigat…

2024

Massive End-to-end Speech Recognition Models with Time Reduction

NAACL 2024long

We investigate massive end-to-end automatic speech recognition (ASR) models with efficiency improvements achieved by time reduction. The encoders of our models use the neural architecture of Google’s universal speech model (USM), with additional funnel pooling layers to significantly reduce the fram…

Cited by 2SourcePDFScholar
2022

Improving Confidence Estimation on Out-of-Domain Data for End-to-End Speech Recognition

ICASSP 2022accepted

As end-to-end automatic speech recognition (ASR) models reach promising performance, various downstream tasks rely on good confidence estimators for these systems. Recent research has shown that model-based confidence estimators have a significant advantage over using the output softmax probabilitie…

Cited by 16SourceScholar
2022

Knowledge Distillation for Neural Transducers from Large Self-Supervised Pre-Trained Models

ICASSP 2022accepted

Self-supervised pre-training is an effective approach to leveraging a large amount of unlabelled data to reduce word error rates (WERs) of automatic speech recognition (ASR) systems. Since it is impractical to use large pre-trained models for many real-world ASR applications, it is desirable to have…

Cited by 29SourceScholar
2021

Confidence Estimation for Attention-Based Sequence-to-Sequence Models for Speech Recognition

ICASSP 2021accepted

For various speech-related tasks, confidence scores from a speech recogniser are a useful measure to assess the quality of transcriptions. In traditional hidden Markov model-based automatic speech recognition (ASR) systems, confidence scores can be reliably obtained from word posteriors in decoding…

Cited by 0SourceScholar
2021

Learning Word-Level Confidence for Subword End-To-End ASR

ICASSP 2021accepted

We study the problem of word-level confidence estimation in subword-based end-to-end (E2E) models for automatic speech recognition (ASR). Although prior works have proposed training auxiliary confidence models for ASR systems, they do not extend naturally to systems that operate on word-pieces (WP)…

Cited by 0SourceScholar
2019

Bi-directional Lattice Recurrent Neural Networks for Confidence Estimation

ICASSP 2019accepted

The standard approach to mitigate errors made by an automatic speech recognition system is to use confidence scores associated with each predicted word. In the simplest case, these scores are word posterior probabilities whilst more complex schemes utilise bi-directional recurrent neural network (Bi…

Cited by 0SourceScholar
2017

Generative Modeling of Audible Shapes for Object Perception

ICCV 2017poster

Humans infer rich knowledge of objects from both auditory and visual cues. Building a machine of such competency, however, is very challenging, due to the great difficulty in capturing large-scale, clean data of objects with both their appearance and the sound they make. In this paper, we present a…

Cited by 44PDFScholar
2017

Shape and Material from Sound

NeurIPS 2017spotlight

Hearing an object falling onto the ground, humans can recover rich information including its rough shape, material, and falling height. In this paper, we build machines to approximate such competency. We first mimic human knowledge of the physical world by building an efficient, physics-based simula…

Cited by 36SourcePDFScholar