FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training
Jiale Huang, Dehong Gao, Jinxia Zhang, Zechao Zhan, Yang Hu, Xin Wang
Abstract
Large-scale Vision-Language Pre-training (VLP) has demonstrated remarkable success in the general domain. However, in the fashion domain, items are distinguished by fine-grained attributes such as texture and material, which are crucial for tasks such as retrieval. Existing models often fail to take advantage of these fine-grained attributes from both text and image modalities. To address the above issue, we propose a novel approach for the fashion domain, Fine-grained Attributes Enhanced VLP (FashionFAE), which focuses on the detailed characteristics of the fashion data. An attribute-emphasized text prediction task is proposed to predict fine-grained attributes of the items. This forces the model to focus on the salient attributes from the text modality. In addition, a novel attribute-promoted image reconstruction task is proposed, which further enhances the fine-grained ability of the model by leveraging the representative attributes from the image modality. Extensive experiments show that FashionFAE outperforms State-Of-The-Art (SOTA) methods, achieving 2.9% and 5.2% improvements in retrieval on sub-test set and full test set, respectively, and an average improvement of 1.6% in recognition tasks.
BibTeX
@inproceedings{icassp2025_fashionfaefinegr,
title = {FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training},
author = {Jiale Huang and Dehong Gao and Jinxia Zhang and Zechao Zhan and Yang Hu and Xin Wang},
booktitle = {ICASSP 2025},
year = {2025}
}