← Search

Amaia Salvador

5 accepted papers

2021

Revamping Cross-Modal Recipe Retrieval With Hierarchical Transformers and Self-Supervised Learning

CVPR 2021poster

Cross-modal recipe retrieval has recently gained substantial attention due to the importance of food in people's lives, as well as the availability of vast amounts of digital cooking recipes and food images to train machine learning models. In this work, we revisit existing approaches for cross-moda…

Cited by 88PDFcodeScholar
2019

Inverse Cooking: Recipe Generation From Food Images

CVPR 2019poster

People enjoy food photography because they appreciate food. Behind each meal there is a story described in a complex recipe and, unfortunately, by simply looking at a food image we do not have access to its preparation process. Therefore, in this paper we introduce an inverse cooking system that rec…

Cited by 210PDFcodeScholar
2019

RVOS: End-To-End Recurrent Network for Video Object Segmentation

CVPR 2019poster

Multiple object video object segmentation is a challenging task, specially for the zero-shot case, when no object mask is given at the initial frame and the model has to find the objects to be segmented along the sequence. In our work, we propose a Recurrent network for multiple object Video Object…

Cited by 285PDFcodeScholar
2019

Wav2Pix: Speech-conditioned Face Generation Using Generative Adversarial Networks

ICASSP 2019accepted

Speech is a rich biometric signal that contains information about the identity, gender and emotional state of the speaker. In this work, we explore its potential to generate face images of a speaker by conditioning a Generative Adversarial Network (GAN) with raw speech input. We propose a deep neura…

Cited by 0SourceScholar
2017

Learning Cross-Modal Embeddings for Cooking Recipes and Food Images

CVPR 2017poster

In this paper, we introduce Recipe1M, a new large-scale, structured corpus of over 1m cooking recipes and 800k food images. As the largest publicly available collection of recipe data, Recipe1M affords the ability to train high-capacity models on aligned, multi-modal data. Accordingly, we train a ne…

Cited by 758PDFScholar