← Search

Zhihong Lei

6 accepted papers

2026

VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation

ICLR 2026poster

Emotion recognition in speech presents a complex multimodal challenge, requiring comprehension of both linguistic content and vocal expressivity, particularly prosodic features such as fundamental frequency, intensity, and temporal dynamics. Although large language models (LLMs) have shown promise i…

Cited by 0SourceScholar
2025

Contextualization of ASR with LLM using phonetic retrieval-based augmentation

ICASSP 2025accepted

Large language models (LLMs) have shown superb capability of modeling multimodal signals including audio and text, allowing the model to generate spoken or textual response given a speech input. However, it remains a challenge for the model to recognize personal named entities, such as contacts in a…

Cited by 0SourceScholar
2024

Conformer-Based Speech Recognition On Extreme Edge-Computing Devices

NAACL 2024industry

With increasingly more powerful compute capabilities and resources in today’s devices, traditionally compute-intensive automatic speech recognition (ASR) has been moving from the cloud to devices to better protect user privacy. However, it is still challenging to implement on-device ASR on resource-…

Cited by 4SourcePDFScholar
2024

Personalization of CTC-Based End-to-End Speech Recognition Using Pronunciation-Driven Subword Tokenization

ICASSP 2024accepted

Recent advances in deep learning and automatic speech recognition have improved the accuracy of end-to-end speech recognition systems, but recognition of personal content such as contact names remains a challenge. In this work, we describe our personalization solution for an end-to-end speech recogn…

Cited by 0SourceScholar
2020

Neural Language Modeling for Named Entity Recognition

COLING 2020main

Named entity recognition is a key component in various natural language processing systems, and neural architectures provide significant improvements over conventional approaches. Regardless of different word embedding and hidden layer structures of the networks, a conditional random field layer is…

2018

Prediction of LSTM-RNN Full Context States as a Subtask for N-Gram Feedforward Language Models

ICASSP 2018accepted

Long short-term memory (LSTM) recurrent neural network language models compress the full context of variable lengths into a fixed size vector. In this work, we investigate the task of predicting the LSTM hidden representation of the full context from a truncated n-gram context as a subtask for train…

Cited by 0SourceScholar