2026
Enhancing Action and Ingredient Modeling for Semantically Grounded Recipe Generation
ICASSP 2026poster
Recent advances in Multimodal Large Language Models (MLMMs) have enabled recipe generation from food images, yet outputs often contain semantically incorrect actions or ingredients despite high lexical scores (e.g., BLEU, ROUGE). To address this gap, we propose a semantically grounded framework that…