ICASSP 2024accepted0 citations

Cooking-Clip: Context-Aware Language-Image Pretraining for Zero-Shot Recipe Generation

Lin Wang, Haithm M. Al-Gunid, Ammar Hawbani, Yan Xiong

Abstract

Cooking is one of the oldest and the most common human activities in everyone’s daily life. Instructional cooking videos have also become one of the most common data sources for multimodal visual understanding researches. Compared to other domains, multimodal cooking videos: 1. not only have significantly stronger cross-modal dependencies between the speech transcriptions and their semantically-aligned visual frames at static time stamps; 2. but also have significantly stronger cross-context dependencies among the sequential steps along the temporal dimension, resulting as an ideal domain for contextualized semantic understanding. We propose Cooking-CLIP, which introduces the concept of language-image pretraining (CLIP) from a general-purpose multimodal embedding problem into a customized recipe generation application. We also propose a context-aware pretraining approach, to facilitate a better CLIP customization to cooking-related applications. Our approach achieves higher text generation accuracies than a strong zero-shot baseline on two instructional cooking video data sets, CrossTask and YouCook2. We also achieve comparative accuracies against a fully-supervised approach, with only a narrow difference, in spite of our zero-shot setting.

BibTeX
@inproceedings{icassp2024_cookingclipconte,
  title = {Cooking-Clip: Context-Aware Language-Image Pretraining for Zero-Shot Recipe Generation},
  author = {Lin Wang and Haithm M. Al-Gunid and Ammar Hawbani and Yan Xiong},
  booktitle = {ICASSP 2024},
  year = {2024}
}