← Search

Xingyu Na

2 accepted papers

2025

Contextualization of ASR with LLM using phonetic retrieval-based augmentation

ICASSP 2025accepted

Large language models (LLMs) have shown superb capability of modeling multimodal signals including audio and text, allowing the model to generate spoken or textual response given a speech input. However, it remains a challenge for the model to recognize personal named entities, such as contacts in a…

Cited by 0SourceScholar
2016

An empirical exploration of CTC acoustic models

ICASSP 2016accepted

The connectionist temporal classification (CTC) loss function has several interesting properties relevant for automatic speech recognition (ASR): applied on top of deep recurrent neural networks (RNNs), CTC learns the alignments between speech frames and label sequences automatically, which removes…

Cited by 0SourceScholar