ICASSP 2025accepted0 citations

Multimodal and Multiple Prompts for Biology

Chonghuinan Wang, Xi Chen, Hongxun Yao

Abstract

Prompt learning for vision-language models, e.g. CLIP, is a rapidly advancing topic, focusing on leveraging pre-trained foundational models to generalize to downstream tasks with minimal training samples. Current prompt-tuning methods primarily optimize on general datasets, yet their application in biological datasets remains limited due to the field’s specificity. To address this gap, we introduce Multimodal and Multiple Prompts for Biology (MMPB). Our approach incorporates a Vision Prompt Injection module, which injects diverse biological visual information from the vision branch into the language branch. Additionally, we propose a multiple prompts scheme tailored to the fine-grained naming conventions in biology. We assess the method’s effectiveness on generation to novel classes, and results demonstrate that MMPB significantly outperforms existing approaches across eight biological datasets. Specifically, MMPB achieves a 14.90% absolute improvement in novel classes performance and an 11.71% gain in harmonic mean compared to the state-of-the-art.

BibTeX
@inproceedings{icassp2025_multimodalandmul,
  title = {Multimodal and Multiple Prompts for Biology},
  author = {Chonghuinan Wang and Xi Chen and Hongxun Yao},
  booktitle = {ICASSP 2025},
  year = {2025}
}