← Search

Zhehuai Chen

20 accepted papers

2025

Anticipating Future with Large Language Model for Simultaneous Machine Translation

NAACL 2025long

Simultaneous machine translation (SMT) takes streaming input utterances and incrementally produces target text. Existing SMT methods only use the partial utterance that has already arrived at the input and the generated hypothesis. Motivated by human interpreters’ technique to forecast future words…

Cited by 0SourcePDFScholar
2025

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators

ICLR 2025poster

An ideal multimodal agent should be aware of the quality of its input modalities. Recent advances have enabled large language models (LLMs) to incorporate auditory systems for handling various speech-related tasks. However, most audio LLMs remain unaware of the quality of the speech they process. Th…

Cited by 1SourcePDFScholar
2025

Chain-of-Thought Prompting for Speech Translation

ICASSP 2025accepted

Large language models (LLMs) have demonstrated remarkable advancements in language understanding and generation. Building on the success of text-based LLMs, recent research has adapted these models to use speech embeddings for prompting, resulting in Speech-LLM models that exhibit strong performance…

Cited by 0SourceScholar
2025

Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

ICASSP 2025accepted

Recent end-to-end speech language models (SLMs) have expanded upon the capabilities of large language models (LLMs) by incorporating pre-trained speech models. However, these SLMs often undergo extensive speech instruction-tuning to bridge the gap between speech and text modalities. This requires si…

Cited by 0SourceScholar
2025

EMMeTT: Efficient Multimodal Machine Translation Training

ICASSP 2025accepted

A rising interest in the modality extension of foundation language models warrants discussion on the most effective, and efficient, multimodal training approach. This work focuses on neural machine translation (NMT) and proposes a joint multimodal training regime of Speech-LLM to include automatic s…

Cited by 0SourceScholar
2025

SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models

ACL 2025long

We introduce Speech-based Intelligence Quotient (SIQ) as a new form of human cognition-inspired evaluation pipeline for voice understanding large language models (LLM_Voice), designed to assess their voice understanding ability. Moving beyond popular voice understanding metrics such as word error ra…

2025

VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning

NAACL 2025long

Recent studies have augmented large language models (LLMs) with speech capabilities, leading to the development of speech language models (SpeechLMs). Earlier SpeechLMs focused on single-turn speech-based question answering (QA), where user input comprised a speech context and a text question. More…

2024

GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators

ACL 2024long

Recent advances in large language models (LLMs) have stepped forward the development of multilingual speech and machine translation by its reduced representation errors and incorporated external knowledge. However, both translation tasks typically utilize beam search decoding and top-1 hypothesis se…

2024

SALM: Speech-Augmented Language Model with in-Context Learning for Speech Recognition and Translation

ICASSP 2024accepted

We present a novel Speech Augmented Language Model (SALM) with multitask and in-context learning capabilities. SALM comprises a frozen text LLM, a audio encoder, a modality adapter module, and LoRA layers to accommodate speech input and associated task instructions. The unified SALM not only achieve…

Cited by 0SourceScholar
2024

Transducers with Pronunciation-Aware Embeddings for Automatic Speech Recognition

ICASSP 2024accepted

This paper proposes Transducers with Pronunciation-aware Embeddings (PET). Unlike conventional Transducers where the decoder embeddings for different tokens are trained independently, the PET model’s decoder embedding incorporates shared components for text tokens with the same or similar pronunciat…

Cited by 0SourceScholar
2023

Accelerating RNN-T Training and Inference Using CTC Guidance

ICASSP 2023accepted

We propose a novel method to accelerate training and inference process of recurrent neural network transducer (RNN-T) based on the guidance from a co-trained connectionist temporal classification (CTC) model. We made a key assumption that if an encoder embedding frame is classified as a blank frame…

Cited by 0SourceScholar
2023

Understanding Shared Speech-Text Representations

ICASSP 2023accepted

Recently, a number of approaches to train speech models by incorporating text into end-to-end models have been developed, with Maestro advancing state-of-the-art automatic speech recognition (ASR) and Speech Translation (ST) performance. In this paper, we expand our understanding of the resulting sh…

Cited by 0SourceScholar
2023

Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-to-Speech

ICASSP 2023accepted

This paper proposes Virtuoso, a massively multilingual speech–text joint semi-supervised learning framework for text-to-speech synthesis (TTS) models. Existing multilingual TTS typically supports tens of languages, which are a small fraction of the thousands of languages in the world. One difficulty…

Cited by 0SourceScholar
2022

Tts4pretrain 2.0: Advancing the use of Text and Speech in ASR Pretraining with Consistency and Contrastive Losses

ICASSP 2022accepted

An effective way to learn representations from untranscribed speech and unspoken text with linguistic/lexical representations derived from synthesized speech was introduced in tts4pretrain [1]. However, the representations learned from synthesized and real speech are likely to be different, potentia…

Cited by 0SourceScholar
2021

An Asynchronous WFST-Based Decoder for Automatic Speech Recognition

ICASSP 2021accepted

We introduce asynchronous dynamic decoder, which adopts an efficient A* algorithm to incorporate big language models in the one-pass decoding for large vocabulary continuous speech recognition. Unlike standard one-pass decoding with on-the-fly composition decoder which might induce a significant com…

Cited by 0SourceScholar
2020

Improving Speech Recognition Using Consistent Predictions on Synthesized Speech

ICASSP 2020accepted

Speech synthesis has advanced to the point of being close to indistinguishable from human speech. However, efforts to train speech recognition systems on synthesized utterances have not been able to show that synthesized data can be effectively used to augment or replace human speech. In this work,…

Cited by 0SourceScholar
2019

End-to-end Contextual Speech Recognition Using Class Language Models and a Token Passing Decoder

ICASSP 2019accepted

End-to-end modeling (E2E) of automatic speech recognition (ASR) blends all the components of a traditional speech recognition system into a single, unified model. Although it simplifies the ASR systems, the unified model is hard to adapt when training and testing data mismatches. In this work, we fo…

Cited by 0SourceScholar