← Search

Yicheng Zhong

3 accepted papers

2026

HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech Understanding

AAAI 2026technical

Recent advances in Speech Large Language Models (Speech LLMs) have led to great progress in speech understanding tasks such as Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER). However, whether these models can achieve human-level auditory perception, particularly in terms of

Cited by 0SourcePDFScholar
2024

ExpCLIP: Bridging Text and Facial Expressions via Semantic Alignment

AAAI 2024technical

The objective of stylized speech-driven facial animation is to create animations that encapsulate specific emotional expressions. Existing methods often depend on pre-established emotional labels or facial expression templates, which may limit the necessary flexibility for accurately conveying user…

Cited by 7SourcePDFScholar
2023

Semi-supervised Speech-driven 3D Facial Animation via Cross-modal Encoding

ICCV 2023poster

Existing Speech-driven 3D facial animation methods typically follow the supervised paradigm, involving regression from speech to 3D facial animation. This paradigm faces two major challenges: the high cost of supervision acquisition, and the ambiguity in mapping between speech and lip movements. To…

Cited by 1PDFScholar