← Search

Ye Gao

7 accepted papers

2026

EMO-TTA: IMPROVING TEST-TIME ADAPTATION OF AUDIO-LANGUAGE MODELS FOR SPEECH EMOTION RECOGNITION

ICASSP 2026oral

Speech emotion recognition (SER) with audio-language models (ALMs) remains vulnerable to distribution shifts at test time, leading to performance degradation in out-of-domain scenarios. Test-time adaptation (TTA) provides a promising solution but often relies on gradient-based updates or prompt tuni…

Cited by 0SourcePDFScholar
2026

EMOTION-ALIGNED GENERATION IN DIFFUSION TEXT TO SPEECH MODELS VIA PREFERENCE-GUIDED OPTIMIZATION

ICASSP 2026oral

Emotional text-to-speech seeks to convey affect while preserving intelligibility and prosody, yet existing methods rely on coarse labels or proxy classifiers and receive only utterance-level feedback. We introduce Emotion-Aware Stepwise Preference Optimization (EASPO), a post-training framework that…

Cited by 0SourcePDFScholar
2026

PLUG-AND-PLAY EMOTION GRAPHS FOR COMPOSITIONAL PROMPTING IN ZERO-SHOT SPEECH EMOTION RECOGNITION

ICASSP 2026poster

Large audio-language models (LALMs) exhibit strong zero-shot performance across speech tasks but struggle with speech emotion recognition (SER) due to weak paralinguistic modeling and limited cross-modal reasoning. We propose Compositional Chain-of-Thought Prompting for Emotion Reasoning (CCoT-Emo),…

Cited by 0SourcePDFScholar
2025

Motion-Feat: Motion Blur-Aware Local Feature Description for Image Matching

IROS 2025

Local feature description is crucial for robotic tasks, yet existing methods struggle with motion blur, a prevalent challenge in high-dynamic and low-light environments. While effective on sharp images, they suffer significant degradation under blur. To address this issue, we propose Motion-Feat, an

Cited by 0SourcecodeScholar
2022

Efficient Video Deblurring Guided by Motion Magnitude

ECCV 2022poster

"Video deblurring is a highly under-constrained problem due to the spatially and temporally varying blur. An intuitive approach for video deblurring includes two steps: a) detecting the blurry region in the current frame; b) utilizing the information from clear regions in adjacent frames for current…

2020

Efficient Spatio-Temporal Recurrent Neural Network for Video Deblurring

ECCV 2020poster

Real-time video deblurring still remains a challenging task due to the complexity of spatially and temporally varying blur itself and the requirement of low computational cost. To improve the network efficiency, we adopt residual dense blocks into RNN cells, so as to efficiently extract the spatial…

Cited by 175SourcePDFScholar
2020

Stereoscopic Flash and No-Flash Photography for Shape and Albedo Recovery

CVPR 2020poster

We present a minimal imaging setup that harnesses both geometric and photometric approaches for shape and albedo recovery. We adopt a stereo camera and a flashlight to capture a stereo image pair and a flash/no-flash pair. From the stereo image pair, we recover a rough shape that captures low-freque…

Cited by 12PDFScholar