ICASSP 2025accepted0 citations

DLM-VMTL: A Double LayerMapper For Heterogeneous Data Video Multi-Task Prompt Learning

Zeyi Bo, Ye Jin, Wuxi Sun

Abstract

In recent years, the parameters of backbones of Video Understanding tasks continue to increase and even reach billion-level. Whether fine-tuning a specific task on the Video Foundation Models(VFMs) or pre-training the model designed for the specific task, incurs significant overhead. How to enable these models to play roles other than those corresponding to their own tasks becomes a worthy issue. Multi-Task Learning(MTL) makes a visual task acquire the rich shareable knowledge from other tasks while joint training. It is fully explored in Image Recognition tasks especially dense predict tasks. Nevertheless, it is rarely used in video domain due to the lack of multi-labels video data. In this paper, a heterogeneous data video multi-task prompt learning (VMTL) method is proposed to address above problem. It’s different from it in image domain, a Double-Layers Mapper(DLM) is proposed to extract the shareable knowledge into visual prompts and align it with representation of primary task to fine-tune the primary task. Extensive experiments prove that our DLM-VMTL performs better than baselines on 6 different video understanding tasks and 11 datasets.

BibTeX
@inproceedings{icassp2025_dlmvmtladoublela,
  title = {DLM-VMTL: A Double LayerMapper For Heterogeneous Data Video Multi-Task Prompt Learning},
  author = {Zeyi Bo and Ye Jin and Wuxi Sun},
  booktitle = {ICASSP 2025},
  year = {2025}
}