Multi-Modal Food Classification in a Diet Tracking System with Spoken and Visual Inputs
Shivani Gowda, Yifan Hu, Mandy Korpusik
Abstract
In this paper, we present multi-modal approaches to diet tracking. As health and well-being become increasingly important, mobile applications for diet tracking attract much interest. However, these applications often require users to log their meals based on relatively unreliable memory recall, thereby underestimating nutritional intake and, thus, undermining the efforts of nutrition tracking. To accurately record dietary intake, there is an increasing need for image computational methods. We investigated multi-modal transfer learning approaches on a novel, food-specific image-text dataset, specifically a Vision-and-Language Transformer that achieves a held-out test set Micro-F1 score of 77.70% and Macro-F1 score of 51.43% for 696 food categories. We aim to give other researchers new insight into the process of developing domain-specific, multi-modal deep learning models with small datasets.
BibTeX
@inproceedings{icassp2023_multimodalfoodcl,
title = {Multi-Modal Food Classification in a Diet Tracking System with Spoken and Visual Inputs},
author = {Shivani Gowda and Yifan Hu and Mandy Korpusik},
booktitle = {ICASSP 2023},
year = {2023}
}