AAAI 2023technical15 citations

DE-net: Dynamic Text-Guided Image Editing Adversarial Networks

Ming Tao, Bing-Kun Bao, Hao Tang, Fei Wu, Longhui Wei, Qi Tian

Abstract

Text-guided image editing models have shown remarkable results. However, there remain two problems. First, they employ fixed manipulation modules for various editing requirements (e.g., color changing, texture changing, content adding and removing), which results in over-editing or insufficient editing. Second, they do not clearly distinguish between text-required and text-irrelevant parts, which leads to inaccurate editing. To solve these limitations, we propose: (i) a Dynamic Editing Block (DEBlock) that composes different editing modules dynamically for various editing requirements. (ii) a Composition Predictor (Comp-Pred), which predicts the composition weights for DEBlock according to the inference on target texts and source images. (iii) a Dynamic text-adaptive Convolution Block (DCBlock) that queries source image features to distinguish text-required parts and text-irrelevant parts. Extensive experiments demonstrate that our DE-Net achieves excellent performance and manipulates source images more correctly and accurately.

BibTeX
@article{Tao_Bao_Tang_Wu_Wei_Tian_2023, title={DE-net: Dynamic Text-Guided Image Editing Adversarial Networks}, volume={37}, url={https://ojs.aaai.org/index.php/AAAI/article/view/26189}, DOI={10.1609/aaai.v37i8.26189}, abstractNote={Text-guided image editing models have shown remarkable results. However, there remain two problems. First, they employ fixed manipulation modules for various editing requirements (e.g., color changing, texture changing, content adding and removing), which results in over-editing or insufficient editing. Second, they do not clearly distinguish between text-required and text-irrelevant parts, which leads to inaccurate editing.
To solve these limitations, we propose:
(i) a Dynamic Editing Block (DEBlock) that composes different editing modules dynamically for various editing requirements.
(ii) a Composition Predictor (Comp-Pred), which predicts the composition weights for DEBlock according to the inference on target texts and source images.
(iii) a Dynamic text-adaptive Convolution Block (DCBlock) that queries source image features to distinguish text-required parts and text-irrelevant parts.
Extensive experiments demonstrate that our DE-Net achieves excellent performance and manipulates source images more correctly and accurately.}, number={8}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Tao, Ming and Bao, Bing-Kun and Tang, Hao and Wu, Fei and Wei, Longhui and Tian, Qi}, year={2023}, month={Jun.}, pages={9971-9979} }
DE-net: Dynamic Text-Guided Image Editing Adversarial Networks · AAAI 2023