← Search

Jianrong Wang

10 accepted papers

2025

A Chinese Expressive Long-dialogue Speech Dataset with Scripts

ICASSP 2025accepted

With the advancement of large-scale models, the demand for emotionally rich, long-context, and highly natural communication in human-computer interaction increases. However, the exploration of long-context or script-level speech conversation tasks remains limited due to the lack of specific supervis…

Cited by 0SourceScholar
2025

Feature-Structure Adaptive Completion Graph Neural Network for Cold-start Recommendation

AAAI 2025technical

The cold-start recommendation has been challenging due to the limited historical interactions for new users and new items. Recently, methods based on meta-learning and graph neural networks have been effective in this problem. However, these methods mainly focus on the missing user-item interactions…

Cited by 0SourcePDFScholar
2025

TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators

ACL 2025finding

Triton, a high-level Python-like language designed for building efficient GPU kernels, is widely adopted in deep learning frameworks due to its portability, flexibility, and accessibility. However, programming and parallel optimization still require considerable trial and error from Triton developer…

2023

Memory-Augmented Contrastive Learning for Talking Head Generation

ICASSP 2023accepted

Given one reference facial image and a piece of speech as input, talking head generation aims to synthesize a realistic-looking talking head video. However, generating a lip-synchronized video with natural head movements is challenging. The same speech clip can generate multiple possible lip and hea…

Cited by 0SourceScholar
2023

Two-Stream Joint-Training for Speaker Independent Acoustic-to-Articulatory Inversion

ICASSP 2023accepted

Acoustic-to-articulatory inversion (AAI) aims to estimate the parameters of articulators from speech audio. There are two common challenges in AAI, which are the limited data and the unsatisfactory performance in speaker independent scenario. Most current works focus on extracting features directly…

Cited by 0SourceScholar
2022

Acoustic-to-Articulatory Inversion Based on Speech Decomposition and Auxiliary Feature

ICASSP 2022accepted

Acoustic-to-articulatory inversion (AAI) is to obtain the movement of articulators from speech signals. Until now, achieving a speaker-independent AAI remains a challenge given the limited data. Besides, most current works only use audio speech as input, causing an inevitable performance bottleneck.…

Cited by 0SourceScholar
2022

Residual-Guided Personalized Speech Synthesis based on Face Image

ICASSP 2022accepted

Previous works derive personalized speech features by training the model on a large dataset composed of his/her audio sounds. It was reported that face information has a strong link with the speech sound. Thus in this work, we innovatively extract personalized speech features from human faces to syn…

Cited by 0SourceScholar
2021

Self-Supervised Depth Estimation Via Implicit Cues from Videos

ICASSP 2021accepted

In self-supervised monocular depth estimation, the depth discontinuity and motion objects' artifacts are still challenging problems. Existing self-supervised methods usually utilize two views to train the depth estimation network and use one single view to make predictions. Compared with static view…

Cited by 0SourceScholar
2016

Continuous ultrasound based tongue movement video synthesis from speech

ICASSP 2016accepted

The movement of tongue plays an important role in pronunciation. Visualizing the movement of tongue can improve speech intelligibility and also helps learning a second language. However, hardly any research has been investigated for this topic. In this paper, a framework to synthesize continuous ult…

Cited by 0SourceScholar