ICASSP 2025accepted0 citations

Learning Hierarchical Attribute Prompt for Vision-Language Models

Jun Liang, Yang Peng, Rui Luo, Yunyu Zou, Yalong Cheng, Bingzhi Chen

Abstract

Prompt learning is a common strategy for adapting Visual Language Models (VLMs) to downstream tasks by fine-tuning prompts for task-specific performance. However, existing methods face two key challenges: overfitting to base classes, which limits generalization to novel classes, and the dependence on manually generated or LLM-based descriptions, which are time-consuming and error-prone. To address these issues, we propose the Learning Hierarchical Attribute Prompt (LHAP) method, which introduces fine-grained semantic alignment through hierarchical prompts. By autonomously extracting visual attributes from images, LHAP generates local-level prompts (LLP) to capture fine-grained semantics and global-level prompts (GLP) to model overall semantics. The combination of LLP and GLP not only improves generalization but also mitigates errors and inefficiencies from manual or LLM-based descriptions. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority and robustness of LHAP over state-of-the-art methods.

BibTeX
@inproceedings{icassp2025_learninghierarch,
  title = {Learning Hierarchical Attribute Prompt for Vision-Language Models},
  author = {Jun Liang and Yang Peng and Rui Luo and Yunyu Zou and Yalong Cheng and Bingzhi Chen},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Learning Hierarchical Attribute Prompt for Vision-Language Models · ICASSP 2025