← Search

Feilong Chen

7 accepted papers

2025

VQTalker: Towards Multilingual Talking Avatars Through Facial Motion Tokenization

AAAI 2025technical

We present VQTalker, a Vector Quantization-based framework for multilingual talking head generation that addresses the challenges of lip synchronization and natural motion across diverse languages. Our approach is grounded in the phonetic principle that human speech comprises a finite set of distinc…

Cited by 0SourcePDFScholar
2024

DiffDub: Person-Generic Visual Dubbing Using Inpainting Renderer with Diffusion Auto-Encoder

ICASSP 2024accepted

Generating high-quality and person-generic visual dubbing remains a challenge. Recent innovation has seen the advent of a two-stage paradigm, decoupling the rendering and lip synchronization process facilitated by intermediate representation as a conduit. Still, previous methodologies rely on rough…

Cited by 0SourceScholar
2024

ViLaS: Exploring the Effects of Vision and Language Context in Automatic Speech Recognition

ICASSP 2024accepted

Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these works have primarily focused on utilizing visual cues derived from human lip motions. In fact, context-dependent visual and…

Cited by 0SourceScholar
2023

DualGATs: Dual Graph Attention Networks for Emotion Recognition in Conversations

ACL 2023long

Capturing complex contextual dependencies plays a vital role in Emotion Recognition in Conversations (ERC). Previous studies have predominantly focused on speaker-aware context modeling, overlooking the discourse structure of the conversation. In this paper, we introduce Dual Graph ATtention network…

2022

A Multi Domain Knowledge Enhanced Matching Network for Response Selection in Retrieval-Based Dialogue Systems

ICASSP 2022accepted

Building a human-machine conversational agent is a core problem in Artificial Intelligence, where knowledge has to be integrated into the model effectively. In this paper, we propose a Multi Domain Knowledge Enhanced Matching Network (MDKEMN) to build retrievalbased dialogue systems that could lever…

Cited by 0SourceScholar
2022

Improving Cross-Modal Understanding in Visual Dialog Via Contrastive Learning

ICASSP 2022accepted

Visual Dialog is a challenging vision-language task since the visual dialog agent needs to answer a series of questions after reasoning over both the image content and dialog history. Though existing methods try to deal with the cross-modal understanding in visual dialog, they are still not enough i…

Cited by 0SourceScholar