← Search

Zhisheng Wang

9 accepted papers

2026

HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech Understanding

AAAI 2026technical

Recent advances in Speech Large Language Models (Speech LLMs) have led to great progress in speech understanding tasks such as Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER). However, whether these models can achieve human-level auditory perception, particularly in terms of

Cited by 0SourcePDFScholar
2026

R4: Nested Reasoning-Retrieval for Reward Modeling in Role-Playing Agents

ICLR 2026poster

Role-playing dialogue presents unique challenges for large language models (LLMs): beyond producing coherent text, models must sustain character persona, integrate contextual knowledge, and convey emotional nuance. Despite strong reasoning abilities, current LLMs often generate dialogue that is lite…

Cited by 0SourceScholar
2025

Eliciting Implicit Acoustic Styles from Open-domain Instructions to Facilitate Fine-grained Controllable Generation of Speech

EMNLP 2025

This paper focuses on generating speech with the acoustic style that meets users’ needs based on their open-domain instructions. To control the style, early work mostly relies on pre-defined rules or templates. The control types and formats are fixed in a closed domain, making it hard to meet divers

Cited by 0SourcePDFScholar
2024

3D Visibility-Aware Generalizable Neural Radiance Fields for Interacting Hands

AAAI 2024technical

Neural radiance fields (NeRFs) are promising 3D representations for scenes, objects, and humans. However, most existing methods require multi-view inputs and per-scene training, which limits their real-life applications. Moreover, current methods focus on single-subject cases, leaving scenes of inte…

2024

ExpCLIP: Bridging Text and Facial Expressions via Semantic Alignment

AAAI 2024technical

The objective of stylized speech-driven facial animation is to create animations that encapsulate specific emotional expressions. Existing methods often depend on pre-established emotional labels or facial expression templates, which may limit the necessary flexibility for accurately conveying user…

Cited by 7SourcePDFScholar
2024

Monocular 3D Hand Mesh Recovery via Dual Noise Estimation

AAAI 2024technical

Current parametric models have made notable progress in 3D hand pose and shape estimation. However, due to the fixed hand topology and complex hand poses, current models are hard to generate meshes that are aligned with the image well. To tackle this issue, we introduce a dual noise estimation metho…

2023

Semi-supervised Speech-driven 3D Facial Animation via Cross-modal Encoding

ICCV 2023poster

Existing Speech-driven 3D facial animation methods typically follow the supervised paradigm, involving regression from speech to 3D facial animation. This paradigm faces two major challenges: the high cost of supervision acquisition, and the ambiguity in mapping between speech and lip movements. To…

Cited by 1PDFScholar
2021

Cross-lingual Text Classification with Heterogeneous Graph Neural Network

ACL 2021short

Cross-lingual text classification aims at training a classifier on the source language and transferring the knowledge to target languages, which is very useful for low-resource languages. Recent multilingual pretrained language models (mPLM) achieve impressive results in cross-lingual classification…