← Search

Yuke Si

4 accepted papers

2026

HarmoniFuse: A Component-Selective and Prompt-Adaptive Framework for Multi-Task Speech Language Modeling

ICASSP 2026poster

Recent advances in large language models have facilitated the development of unified speech language models (SLMs) capable of supporting multiple speech tasks within a shared architecture. However, tasks such as automatic speech recognition (ASR) and speech emotion recognition (SER) rely on distinct…

Cited by 0SourcePDFScholar
2022

Cache: Modeling Contribution-Aware Context Hierarchically for Long-Range Dialogue State Tracking

ICASSP 2022accepted

Recently, many studies on dialogue state tracking (DST) based on the copy-augmented encoder-decoder framework have been proposed and have achieved encouraging performance. However, these studies commonly lose earlier information during encoding the long dialogues with RNNs, and have difficulty for t…

Cited by 0SourceScholar
2022

Speech Emotion Recognition with Co-Attention Based Multi-Level Acoustic Information

ICASSP 2022accepted

Speech Emotion Recognition (SER) aims to help the machine to understand human’s subjective emotion from only audio in-formation. However, extracting and utilizing comprehensive in-depth audio information is still a challenging task. In this paper, we propose an end-to-end speech emotion recognition…

Cited by 0SourceScholar
2020

A Hierarchical Model for Dialog Act Recognition Considering Acoustic and Lexical Context Information

ICASSP 2020accepted

Dialog act recognition (DAR) is important to capture speakers' intention in a dialog system. Traditional methods commonly use the lexical information from transcripts, acoustic information from speech, and dialog context information to do DAR. However, in these methods, textual context information m…

Cited by 0SourceScholar