Learning Structural Representations for Recipe Generation and Food Retrieval
Food is significant to human daily life. In this paper, we are interested in learning structural representations for lengthy recipes, that can benefit the recipe generation and food cross-modal retrieval tasks. Different from the common vision-language data, here the food images contain mixed ingredients and target recipes are lengthy paragraphs, where we do not have annotations on structure information. To address the above limitations, we propose a novel method to unsupervisedly learn the sentence-level tree structures for the cooking recipes. Our approach brings together several novel ideas in a systematic framework: (1) exploiting an unsupervised learning approach to obtain the sentence-level tree structure labels before training; (2) generating trees of target recipes from images with the supervision of tree structure labels learned from (1); and (3) integrating the learned tree structures into the recipe generation and food cross-modal retrieval procedure. Our proposed model can produce good-quality sentence-level tree structures and coherent recipes. We achieve the state-of-the-art recipe generation and food cross-modal retrieval performance on the benchmark Recipe1M dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-Modal RetrievalImage CaptioningRecipe GenerationRetrievalSentenceSimilar Papers 제목 키워드 기반
Cross-Modal Food Retrieval: Learning a Joint Embedding of Food Images and Recipes with Semantic Consistency and Attention Mechanism
Food retrieval is an important task to perform analysis of food-related information, where we are interested in retrieving relevant information about the queried food item such as ingredients, cooking instructions, etc. …
Cross-Modal RetrievalRetrievalMALM: Mask Augmentation based Local Matching for Food-Recipe Retrieval
Image-to-recipe retrieval is a challenging vision-to-language task of significant practical value. The main challenge of the task lies in the ultra-high redundancy in the long recipe and the large variation reflected in …
Image-text matchingRetrievalText MatchingCHEF: Cross-modal Hierarchical Embeddings for Food Domain Retrieval
Despite the abundance of multi-modal data, such as image-text pairs, there has been little effort in understanding the individual entities and their different roles in the construction of these data instances. In this wo…
Cross-Modal RetrievalRetrievalRecipe2Vec: Multi-modal Recipe Representation Learning with Graph Neural Networks
Learning effective recipe representations is essential in food studies. Unlike what has been developed for image-based recipe retrieval or learning structural text embeddings, the combined effect of multi-modal informati…
Adversarial AttackGraph Neural NetworkNode ClassificationRepresentation Learning+1FIRE: Food Image to REcipe generation
Food computing has emerged as a prominent multidisciplinary field of research in recent years. An ambitious goal of food computing is to develop end-to-end intelligent systems capable of autonomously producing recipe inf…
DecoderLanguage ModellingLarge Language ModelRecipe Generation+1