← Search

Xinlu Yu

2 accepted papers

2025

Co-Speech Gesture Video Generation with Implicit Motion-Audio Entanglement

CVPR 2025poster

Co-speech gestures are essential to non-verbal communication, enhancing both the naturalness and effectiveness of human interaction. Although recent methods have made progress in generating co-speech gesture videos, many rely on strong visual controls, such as pose images or TPS keypoint movements,…

2022

Self-supervised Cross-modal Pretraining for Speech Emotion Recognition and Sentiment Analysis

EMNLP 2022finding

Multimodal speech emotion recognition (SER) and sentiment analysis (SA) are important techniques for human-computer interaction. Most existing multimodal approaches utilize either shallow cross-modal fusion of pretrained features, or deep cross-modal fusion with raw features. Recently, attempts have…