← Search

Shangfei Wang

19 accepted papers

2026

ConsistTalk: Intensity Controllable Temporally Consistent Talking Head Generation with Diffusion Noise Search

AAAI 2026technical

Recent advancements in video diffusion models have significantly enhanced audio-driven portrait animation. However, current methods still suffer from flickering, identity drift, and poor audio-visual synchronization. These issues primarily stem from entangled appearance-motion representations and un

Cited by 0SourcePDFScholar
2026

DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement

ICML 2026poster

Unified Multimodal models (UMMs) built on a single architecture have shown impressive performance in both understanding and generation. We identify a fundamental challenge lies in inductive biases induced by distinct supervision signals: generation branch prefers high-fidelity, fine-grained represen…

Cited by 0SourceScholar
2026

Learning Knowledge from Textual Descriptions for 3D Human Pose Estimation

AAAI 2026technical

Mainstream 3D human pose estimation methods directly predict 3D coordinates of joints from 2D keypoints, suffering from severe depth ambiguity. Pose textual descriptions contain abundant semantic information, which facilitates the model to learn the spatial relationship among different body parts, p

Cited by 0SourcePDFScholar
2026

MIRRORTALK: FORGING PERSONALIZED AVATARS VIA DISENTANGLED STYLE AND HIERARCHICAL MOTION CONTROL

ICASSP 2026poster

Synthesizing personalized talking faces that uphold and highlight a speaker's unique style while maintaining lip-sync accuracy remains a significant challenge. A primary limitation of existing approaches is the intrinsic confounding of speaker-specific talking style and semantic content within facia…

Cited by 0SourcePDFScholar
2025

From Traits to Empathy: Personality-Aware Multimodal Empathetic Response Generation

COLING 2025main

Empathetic dialogue systems improve user experience across various domains. Existing approaches mainly focus on acquiring affective and cognitive knowledge from text, but neglect the unique personality traits of individuals and the inherently multimodal nature of human face-to-face conversation. To…

Cited by 2SourcePDFScholar
2025

Integrating Visual Modalities with Large Language Models for Mental Health Support

COLING 2025main

Current work of mental health support primarily utilizes unimodal textual data and often fails to understand and respond to users’ emotional states comprehensively. In this study, we introduce a novel framework that enhances Large Language Model (LLM) performance in mental health dialogue systems by…

Cited by 1SourcePDFScholar
2024

1DFormer: A Transformer Architecture Learning 1D Landmark Representations for Facial Landmark Tracking

IJCAI 2024poster

Recently, heatmap regression methods based on 1D landmark representations have shown prominent performance on locating facial landmarks. However, previous methods ignored to make deep explorations on the good potentials of 1D landmark representations for sequential and structural modeling of multi…

Cited by 0SourcePDFScholar
2024

FreqMark: Invisible Image Watermarking via Frequency Based Optimization in Latent Space

NeurIPS 2024poster

Invisible watermarking is essential for safeguarding digital content, enabling copyright protection and content authentication. However, existing watermarking methods fall short in robustness against regeneration attacks. In this paper, we propose a novel method called FreqMark that involves uncons…

Cited by 0SourcePDFScholar
2017

A Multimodal Deep Regression Bayesian Network for Affective Video Content Analyses

ICCV 2017poster

The inherent dependencies between visual elements and aural elements are crucial for affective video content analyses, yet have not been successfully exploited. Therefore, we propose a multimodal deep regression Bayesian network (MMDRBN) to capture the dependencies between visual elements and aural…

Cited by 24PDFScholar