← Search

Fei Ding

12 accepted papers

2026

Scaling Multi-Identity Consistency for Image Customization via Multi-to-Multi Matching Paradigm

CVPR 2026

Recent advancements in image customization exhibit a wide range of application prospects due to stronger customization capabilities. However, since we humans are more sensitive to faces, a significant challenge remains in preserving consistent identity while avoiding identity confusion with multi-re

Cited by 0SourcecodeScholar
2026

Unified Customized Generation by Disentangled Reward Modeling

CVPR 2026

Existing literature typically treats various customized generation tasks (e.g., subject-customized generation, style-customized generation) as distinct and disjoint problems, with each task focusing solely on customizing a specific aspect of the reference image. However, we argue that the objectives

Cited by 0SourcecodeScholar
2025

Clear Up Confusion: Iterative Differential Generation for Fine-grained Intent Detection with Contrastive Feedback

COLING 2025main

Fine-grained intent detection involves identifying a large number of classes with subtle variations. Recently, generating pseudo samples via large language models has attracted increasing attention to alleviate the data scarcity caused by emerging new intents. However, these methods generate samples…

Cited by 0SourcePDFScholar
2025

Less-to-More Generalization: Unlocking More Controllability by In-Context Generation

ICCV 2025poster

Although subject-driven generation has been extensively explored in image generation due to its wide applications, it still has challenges in data scalability and subject expansibility. For the first challenge, moving from curating single-subject datasets to multiple-subject ones and scaling them is…

2025

Online Adaptive Keypoint Extraction for Visual Odometry Across Different Scenes

RA-L 2025

Visual odometry needs to be robust against various environmental changes. Although Deep learning (DL) based methods can bring more robust features to visual odometry (VO) than traditional methods, the gap between training and test dataset restricts the performance of DL-based methods when encounteri

Cited by 18SourceScholar
2025

UE-Extractor: A Grid-to-Point Ground Extraction Framework for Unstructured Environments Using Adaptive Grid Projection

RA-L 2025

Ground point cloud extraction is crucial for route planning of autonomous vehicles in unstructured environments. However, mainstream point cloud extraction methods are susceptible to inaccuracies due to the indistinct obstacle-ground boundary. Furthermore, addressing uneven feature distribution usua

Cited by 18SourceScholar
2024

From Discrimination to Generation: Low-Resource Intent Detection with Language Model Instruction Tuning

ACL 2024findings

Intent detection aims to identify user goals from utterances, and is a ubiquitous step towards the satisfaction of user desired needs in many interaction systems. As dynamic and varied intents arise, models that are capable of identifying new intents promptly are required. However, existing studies…

Cited by 3SourcePDFScholar
2024

MSSTNet: A Multi-Scale Spatio-Temporal CNN-Transformer Network for Dynamic Facial Expression Recognition

ICASSP 2024accepted

Unlike typical video action recognition, Dynamic Facial Expression Recognition (DFER) does not involve distinct moving targets but relies on localized changes in facial muscles. Addressing this distinctive attribute, we propose a MultiScale Spatio-temporal CNN-Transformer network (MSST-Net). Our app…

Cited by 0SourceScholar
2023

Dual Class Knowledge Propagation Network for Multi-label Few-shot Intent Detection

ACL 2023long

Multi-label intent detection aims to assign multiple labels to utterances and attracts increasing attention as a practical task in task-oriented dialogue systems. As dialogue domains change rapidly and new intents emerge fast, the lack of annotated data motivates multi-label few-shot intent detectio…

Cited by 10SourcePDFScholar
2022

Multi-level Distillation of Semantic Knowledge for Pre-training Multilingual Language Model

EMNLP 2022main

Pre-trained multilingual language models play an important role in cross-lingual natural language understanding tasks. However, existing methods did not focus on learning the semantic structure of representation, and thus could not optimize their performance. In this paper, we propose Multi-level Mu…

Cited by 6SourcePDFScholar
2022

XMP-Font: Self-Supervised Cross-Modality Pre-Training for Few-Shot Font Generation

CVPR 2022poster

Generating a new font library is a very labor-intensive and time-consuming job for glyph-rich scripts. Few-shot font generation is thus required, as it requires only a few glyph references without fine-tuning during test. Existing methods follow the style-content disentanglement paradigm, and expect…

Cited by 58PDFScholar
2021

Multi-Shot Temporal Event Localization: A Benchmark

CVPR 2021poster

Current developments in temporal event or action localization usually target actions captured by a single camera. However, extensive events or actions in the wild may be captured as a sequence of shots by multiple cameras at different positions. In this paper, we propose a new and challenging task c…

Cited by 109PDFcodeScholar