← Search

Jia Jia

27 accepted papers

2026

Emotion-Conditioned Motion Sub-spaces with Flow Matching for Real-Time Audio-Driven Talking Heads

AAAI 2026technical

Recent advances in audio-driven talking-head synthesis have brought lip-sync precision close to human perception, yet emotional fidelity and real-time inference remain open challenges. Existing pipelines typically disentangle lip articulation, facial expression, and head pose in latent space; this

Cited by 0SourcePDFScholar
2026

From Natural Alignment to Conditional Controllability in Multimodal Dialogue

ICLR 2026poster

The recent advancement of Artificial Intelligence Generated Content (AIGC) has led to significant strides in modeling human interaction, particularly in the context of multi-modal dialogue. While current methods impressively generates realistic dialogue in speech and vision modalities, challenges r…

Cited by 0SourcecodeScholar
2026

VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive Generation

CVPR 2026

Visual generation is dominated by three paradigms: AutoRegressive (AR), diffusion, and Visual AutoRegressive (VAR) models. Unlike AR and diffusion, VARs operate on heterogeneous input structures across their generation steps, which creates severe asynchronous policy conflicts. This issue becomes par

Cited by 0SourcecodeScholar
2026

Value-as-Return: A Two-Stage Framework to Align on the Optimal Score Function

ICML 2026poster

Reinforcement learning with diffusion models has shown strong potential, but existing approaches such as variants of Direct Preference Optimization (DPO) often rely on an inaccurate simplification: they equate trajectory likelihoods with final-state probabilities. This mismatch leads to suboptimal a…

Cited by 0SourceScholar
2025

Minimal Impact ControlNet: Advancing Multi-ControlNet Integration

ICLR 2025poster

With the advancement of diffusion models, there is a growing demand for high-quality, controllable image generation, particularly through methods that utilize one or multiple control signals based on ControlNet. However, in current ControlNet training, each control is designed to influence all areas…

Cited by 0SourcePDFScholar
2024

DanceCamera3D: 3D Camera Movement Synthesis with Music and Dance

CVPR 2024poster

Choreographers determine what the dances look like while cameramen determine the final presentation of dances. Recently various methods and datasets have showcased the feasibility of dance synthesis. However camera movement synthesis with music and dance remains an unsolved challenging problem due t…

2024

Inner Classifier-Free Guidance and Its Taylor Expansion for Diffusion Models

ICLR 2024poster

Classifier-free guidance (CFG) is a pivotal technique for balancing the diversity and fidelity of samples in conditional diffusion models. This approach involves utilizing a single model to jointly optimize the conditional score predictor and unconditional score predictor, eliminating the need for a…

Cited by 2SourcePDFScholar
2024

Skinned Motion Retargeting with Dense Geometric Interaction Perception

NeurIPS 2024spotlight

Capturing and maintaining geometric interactions among different body parts is crucial for successful motion retargeting in skinned characters. Existing approaches often overlook body geometries or add a geometry correction stage after skeletal motion retargeting. This results in conflicts between s…

2024

VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding

NeurIPS 2024poster

Existing Video Corpus Moment Retrieval (VCMR) is limited to coarse-grained understanding that hinders precise video moment localization when given fine-grained queries. In this paper, we propose a more challenging fine-grained VCMR benchmark requiring methods to localize the best-matched moment from…

2023

MSNet: A Deep Architecture Using Multi-Sentiment Semantics for Sentiment-Aware Image Style Transfer

ICASSP 2023accepted

Sentiment plays an essential role in people’s perception of images. To incorporate the sentiment information into the image style transfer task for better sentiment-aware performance, we introduce a new task named sentiment-aware image style transfer. To solve this problem, we first introduce a nove…

Cited by 0SourceScholar
2023

SDDM: Score-Decomposed Diffusion Models on Manifolds for Unpaired Image-to-Image Translation

ICML 2023poster

Recent score-based diffusion models (SBDMs) show promising results in unpaired image-to-image translation (I2I). However, existing methods, either energy-based or statistically-based, provide no explicit form of the interfered intermediate generative distributions. This work presents a new score-dec…

Cited by 23SourcePDFScholar
2023

Salient Co-Speech Gesture Synthesizing with Discrete Motion Representation

ICASSP 2023accepted

Synthesizing co-speech gestures is challenging because the mapping from speech to gesticulation is inherently non-deterministic. When giving talks, people conduct not only gentle and rhythmic motions but also abrupt and salient gesticulations. Most previous research efforts, however, ignore this nat…

Cited by 0SourceScholar
2023

What Does Your Face Sound Like? 3D Face Shape towards Voice

AAAI 2023technical

Face-based speech synthesis provides a practical solution to generate voices from human faces. However, directly using 2D face images leads to the problems of uninterpretability and entanglement. In this paper, to address the issues, we introduce 3D face shape which (1) has an anatomical relationshi…

2021

Inferring Emotion from Large-scale Internet Voice Data: A Semi-supervised Curriculum Augmentation based Deep Learning Approach

AAAI 2021technical

Effective emotion inference from user queries helps to give a more personified response for Voice Dialogue Applications(VDAs). The tremendous amounts of VDA users bring in diverse emotion expressions. How to achieve a high emotion inferring performance from large-scale Internet Voice Data in VDAs? T…

Cited by 16SourcePDFScholar
2020

Cross-VAE: Towards Disentangling Expression from Identity For Human Faces

ICASSP 2020accepted

Facial expression and identity are two independent yet intertwined components for representing a face. For facial expression recognition, identity can contaminate the training procedure by providing tangled but irrelevant information. In this paper, we propose to learn clearly disentangled and discr…

Cited by 0SourceScholar
2019

A Compact Framework for Voice Conversion Using Wavenet Conditioned on Phonetic Posteriorgrams

ICASSP 2019accepted

Voice conversion can benefit from WaveNet vocoder with improvement in converted speech's naturalness and quality. However, nowadays approaches segregate the training of conversion module and WaveNet vocoder towards different optimization objectives, which might lead to the difficulty in model tuning…

Cited by 0SourceScholar
2019

Dilated Residual Network with Multi-head Self-attention for Speech Emotion Recognition

ICASSP 2019accepted

Speech emotion recognition (SER) plays an important role in intelligent speech interaction. One vital challenge in SER is to extract emotion-relevant features from speech signals. In state-of-the-art SER techniques, deep learning methods, e.g, Convolutional Neural Networks (CNNs), are widely employe…

Cited by 0SourceScholar
2019

Learning Discriminative Features from Spectrograms Using Center Loss for Speech Emotion Recognition

ICASSP 2019accepted

Identifying the emotional state from speech is essential for the natural interaction of the machine with the speaker. However, extracting effective features for emotion recognition is difficult, as emotions are ambiguous. We propose a novel approach to learn discriminative features from variable len…

Cited by 0SourceScholar
2019

Modality Attention for End-to-end Audio-visual Speech Recognition

ICASSP 2019accepted

Audio-visual speech recognition (AVSR) system is thought to be one of the most promising solutions for robust speech recognition, especially in noisy environment. In this paper, we propose a novel multimodal attention based method for audio-visual speech recognition which could automatically learn t…

Cited by 0SourceScholar
2018

Emphatic Speech Generation with Conditioned Input Layer and Bidirectional LSTMS for Expressive Speech Synthesis

ICASSP 2018accepted

By highlighting the focus of an utterance to draw attention, emphasis in speech interaction plays an important role for speaker intention expressing and understanding. Therefore, emphatic speech synthesis draws increasing interest in the text-to-speech (TTS) area. For emphatic speech synthesis, thre…

Cited by 0SourceScholar
2017

Inferring emotions from heterogeneous social media data: A Cross-media Auto-Encoder solution

ICASSP 2017accepted

Social media is rocking the world in recent year, which makes modeling social media contents important. However, the heterogeneity of social media data is the main constraint. This paper focuses on inferring emotions from large-scale social media data. Tweets on social media platform, always contain…

Cited by 0SourceScholar
2017

Learning cross-lingual knowledge with multilingual BLSTM for emphasis detection with limited training data

ICASSP 2017accepted

Bidirectional long short-term memory (BLSTM) recurrent neural network (RNN) has achieved state-of-the-art performance in many sequence processing problems given its capability in capturing contextual information. However, for languages with limited amount of training data, it is still difficult to o…

Cited by 0SourceScholar
2016

Low level descriptors based DBLSTM bottleneck feature for speech driven talking avatar

ICASSP 2016accepted

Speech is bimodal in nature. There are close correlations between the acoustic speech signals and the visual gestures such as lip movements, facial expressions and head motions. For speech driven talking avatar, how to derive more representative acoustic features from which to predict more accurate…

Cited by 0SourceScholar
2015

HMM-based emphatic speech synthesis for corrective feedback in computer-aided pronunciation training

ICASSP 2015accepted

This paper investigates the incorporation of hidden Markov model (HMM) based emphatic speech synthesis for audio exaggeration into an audio-visual speech synthesis framework for the corrective feedback in computer-aided pronunciation training (CAPT). To improve the voice quality of the synthetic emp…

Cited by 0SourceScholar