← Search

Fei Du

10 accepted papers

2026

LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation

ICLR 2026poster

Recent advances in diffusion models have significantly improved text-to-video generation, enabling personalized content creation with fine-grained control over both foreground and background elements. However, precise face–attribute alignment across subjects remains challenging, as existing methods…

Cited by 0SourcecodeScholar
2025

Over-Generation and Compaction: A Prompting Strategy for Procedural Text Adaptation with Large Language Models

EMNLP 2025

Procedural text adaptation—such as modifying recipes or revising instructional guides—has traditionally relied on specialized models extensively fine‐tuned for specific domains. To address the scalability limitations of such approaches, recent research has increasingly turned to general‐purpose larg

2025

RealisID: Scale-Robust and Fine-Controllable Identity Customization via Local and Global Complementation

AAAI 2025technical

Recently, the success of text-to-image synthesis has greatly advanced the development of identity customization techniques, whose main goal is to produce realistic identity-specific photographs based on text prompts and reference face images. However, it is difficult for existing identity customizat…

2025

STE-Mamba: Automated Multimodal Depression Detection through Emotional Analysis and Spatio-Temporal Information Ensemble

ICASSP 2025accepted

Automatic Depression Detection (ADD) garners widespread attention due to its convenience and objectivity. While existing research makes significant progress, challenges remain. First, most current ADD methods struggle to balance computational overhead and prediction accuracy. Second, these methods p…

Cited by 0SourceScholar
2024

SHMT: Self-supervised Hierarchical Makeup Transfer via Latent Diffusion Models

NeurIPS 2024poster

This paper studies the challenging task of makeup transfer, which aims to apply diverse makeup styles precisely and naturally to a given facial image. Due to the absence of paired data, current methods typically synthesize sub-optimal pseudo ground truths to guide the model training, resulting in l…

2023

Efficient Mask Correction for Click-Based Interactive Image Segmentation

CVPR 2023poster

The goal of click-based interactive image segmentation is to extract target masks with the input of positive/negative clicks. Every time a new click is placed, existing methods run the whole segmentation network to obtain a corrected mask, which is inefficient since several clicks may be needed to r…

2023

Global and Local Mixture Consistency Cumulative Learning for Long-Tailed Visual Recognitions

CVPR 2023poster

In this paper, our goal is to design a simple learning paradigm for long-tail visual recognition, which not only improves the robustness of the feature extractor but also alleviates the bias of the classifier towards head classes while reducing the training skills and overhead. We propose an efficie…

2023

SwinRDM: Integrate SwinRNN with Diffusion Model towards High-Resolution and High-Quality Weather Forecasting

AAAI 2023technical

Data-driven medium-range weather forecasting has attracted much attention in recent years. However, the forecasting accuracy at high resolution is unsatisfactory currently. Pursuing high-resolution and high-quality weather forecasting, we develop a data-driven model SwinRDM which integrates an impro…

Cited by 55SourcePDFScholar