← Search

Zhengcong Fei

12 accepted papers

2024

Prefix-diffusion: A Lightweight Diffusion Model for Diverse Image Captioning

COLING 2024main

While impressive performance has been achieved in image captioning, the limited diversity of the generated captions and the large parameter scale remain major barriers to the real-word application of these systems. In this work, we propose a lightweight image captioning network in combination with c…

2024

Tuning-Free Inversion-Enhanced Control for Consistent Image Editing

AAAI 2024technical

Consistent editing of real images is a challenging task, as it requires performing non-rigid edits (e.g., changing postures) to the main objects in the input image without changing their identity or attributes. To guarantee consistent attributes, some existing methods fine-tune the entire model or t…

Cited by 12SourcePDFScholar
2023

Masked Auto-Encoders Meet Generative Adversarial Networks and Beyond

CVPR 2023poster

Masked Auto-Encoder (MAE) pretraining methods randomly mask image patches and then train a vision Transformer to reconstruct the original pixels based on the unmasked patches. While they demonstrates impressive performance for downstream vision tasks, it generally requires a large amount of training…

Cited by 20SourcePDFScholar
2023

Uncertainty-Aware Image Captioning

AAAI 2023technical

It is well believed that the higher uncertainty in a word of the caption, the more inter-correlated context information is required to determine it. However, current image captioning methods usually consider the generation of all words in a sentence sequentially and equally. In this paper, we propos…

Cited by 19SourcePDFScholar
2022

Selecting Stickers in Open-Domain Dialogue through Multitask Learning

ACL 2022findings

With the increasing popularity of online chatting, stickers are becoming important in our online communication. Selecting appropriate stickers in open-domain dialogue requires a comprehensive understanding of both dialogues and stickers, as well as the relationship between the two types of modalitie…

2021

Conversations Are Not Flat: Modeling the Dynamic Information Flow across Dialogue Utterances

ACL 2021long

Nowadays, open-domain dialogue models can generate acceptable responses according to the historical context based on the large-scale pre-trained language models. However, they generally concatenate the dialogue history directly as the model input to predict the response, which we named as the flat p…

2021

Memory-Augmented Image Captioning

AAAI 2021technical

Current deep learning-based image captioning systems have been proven to store practical knowledge with their parameters and achieve competitive performances in the public datasets. Nevertheless, their ability to access and precisely manipulate the mastered knowledge is still limited. Besides, provi…

Cited by 38SourcePDFScholar