← Search

Qifei Li

8 accepted papers

2026

KPLM-STA: Physically-Accurate Shadow Synthesis for Human Relighting via Keypoint-Based Light Modeling

AAAI 2026technical

Image composition aims to seamlessly integrate a foreground object into a background, where generating realistic and geometrically accurate shadows remains a persistent challenge. While recent diffusion-based methods have outperformed GAN-based approaches, existing techniques, such as the diffusion-

Cited by 0SourcePDFScholar
2025

DetailTTS: Learning Residual Detail Information for Zero-shot Text-to-speech

ICASSP 2025accepted

Traditional text-to-speech (TTS) systems often face challenges in aligning text and speech, leading to the omission of critical linguistic and acoustic details. This misalignment creates an information gap, which existing methods attempt to address by incorporating additional inputs, but these often…

Cited by 0SourceScholar
2025

LoD: Loss-difference OOD Detection by Intentionally Label-Noisifying Unlabeled Wild Data

IJCAI 2025

Using unlabeled wild data containing both in-distribution (ID) and out-of-distribution (OOD) data to improve the safety and reliability of models has recently received increasing attention. Existing methods either design customized losses for labeled ID and unlabeled wild data then perform joint opt

2024

Concss: Contrastive-based Context Comprehension for Dialogue-Appropriate Prosody in Conversational Speech Synthesis

ICASSP 2024accepted

Conversational speech synthesis (CSS) incorporates historical dialogue as supplementary information with the aim of generating speech that has dialogue-appropriate prosody. While previous methods have already delved into enhancing context comprehension, context representation still lacks effective r…

Cited by 0SourceScholar
2024

Frame-Level Emotional State Alignment Method for Speech Emotion Recognition

ICASSP 2024accepted

Speech emotion recognition (SER) systems aim to recognize human emotional state during human-computer interaction. Most existing SER systems are trained based on utterance-level labels. However, not all frames in an audio have affective states consistent with utterance-level label, which makes it di…

Cited by 0SourceScholar
2023

Learning to Predict Persona Information for Dialogue Personalization without Explicit Persona Description

ACL 2023findings

Personalizing dialogue agents is important for dialogue systems to generate more specific,consistent, and engaging responses. However, most current dialogue personalization approaches rely on explicit persona descriptions during inference, which severely restricts its application. In this paper, we…

Cited by 5SourcePDFScholar
2021

Learning from Perturbations: Diverse and Informative Dialogue Generation with Inverse Adversarial Training

ACL 2021long

In this paper, we propose Inverse Adversarial Training (IAT) algorithm for training neural dialogue systems to avoid generic responses and model dialogue history better. In contrast to standard adversarial training algorithms, IAT encourages the model to be sensitive to the perturbation in the dialo…

Cited by 25SourcePDFScholar