← Search

Zhaocheng Huang

7 accepted papers

2023

End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation

EMNLP 2023long main

Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversations by multiple speakers. In this paper, we tackle single-channel multi-speaker conversational ST with an end-to-end an…

Cited by 0SourcecodeScholar
2022

Representation Learning Through Cross-Modal Conditional Teacher-Student Training For Speech Emotion Recognition

ICASSP 2022accepted

Generic pre-trained speech and text representations promise to reduce the need for large labeled datasets on specific speech and language tasks. However, it is not clear how to effectively adapt these representations for speech emotion recognition. Recent public benchmarks show the efficacy of sever…

Cited by 0SourceScholar
2021

Automatic Elicitation Compliance for Short-Duration Speech Based Depression Detection

ICASSP 2021accepted

Detecting depression from the voice in naturalistic environments is challenging, particularly for short-duration audio recordings. This enhances the need to interpret and make optimal use of elicited speech. The rapid consonant-vowel syllable combination ‘pataka’ has frequently been selected as a cl…

Cited by 0SourceScholar
2020

Exploiting Vocal Tract Coordination Using Dilated CNNS For Depression Detection In Naturalistic Environments

ICASSP 2020accepted

Depression detection from speech continues to attract significant research attention but remains a major challenge, particularly when the speech is acquired from diverse smartphones in natural environments. Analysis methods based on vocal tract coordination have shown great promise in depression and…

Cited by 0SourceScholar
2019

Speech Landmark Bigrams for Depression Detection from Naturalistic Smartphone Speech

ICASSP 2019accepted

Detection of depression from speech has attracted significant research attention in recent years but remains a challenge, particularly for speech from diverse smartphones in natural environments. This paper proposes two sets of novel features based on speech landmark bigrams associated with abrupt s…

Cited by 0SourceScholar
2017

A PLLR and multi-stage Staircase Regression framework for speech-based emotion prediction

ICASSP 2017accepted

Continuous prediction of dimensional emotions (e.g. arousal and valence) has attracted increasing research interest recently. When processing emotional speech signals, phonetic features have been rarely used due to the assumption that phonetic variability is a confounding factor that degrades emotio…

Cited by 0SourceScholar