paper-with-me

Papers

Visual Recipe Flow: A Dataset for Learning Visual State Changes of Objects with Recipe Flows

2022-09-13 · COLING 2022 10 · Keisuke Shirai, Atsushi Hashimoto, Taichi Nishimura, Hirotaka Kameko, Shuhei Kurita, Yoshitaka Ushiku, Shinsuke Mori

We present a new multimodal dataset called Visual Recipe Flow, which enables us to learn each cooking action result in a recipe text. The dataset consists of object state changes and the workflow of the recipe text. The state change is represented as an image pair, while the workflow is represented as a recipe flow graph (r-FG). The image pairs are grounded in the r-FG, which provides the cross-modal relation. With our dataset, one can try a range of applications, from multimodal commonsense reasoning and procedural text generation.

📄 PDF Abstract BibTeX arXiv:2209.05840

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

Visual Grounding Annotation of Recipe Flow Graph

2020-05-01 · LREC 2020 5 · Taichi Nishimura, Suzushi Tomori, Hayato Hashimoto, Atsushi Hashimoto 외

In this paper, we provide a dataset that gives visual grounding annotations to recipe flow graphs. A recipe flow graph is a representation of the cooking workflow, which is designed with the aim of understanding the work…

Visual Grounding

Multi-modal Cooking Workflow Construction for Food Recipes

2020-08-20 · Liangming Pan, Jingjing Chen, Jianlong Wu, Shaoteng Liu 외

Understanding food recipe requires anticipating the implicit causal effects of cooking actions, such that the recipe can be converted into a graph describing the temporal workflow of the recipe. This is a non-trivial tas…

Common Sense ReasoningDecoder

OSCAR: Object Status and Contextual Awareness for Recipes to Support Non-Visual Cooking

2025-03-07 · Franklin Mingzhe Li, Kaitlyn Ng, Bin Zhu, Patrick Carrington

Following recipes while cooking is an important but difficult task for visually impaired individuals. We developed OSCAR (Object Status Context Awareness for Recipes), a novel approach that provides recipe progress track…

Object

TextLDM: Language Modeling with Continuous Latent Diffusion

2026-05-08 · Jiaxiu Jiang, Jingjing Ren, Wenbo Li, Bo Wang 외 arxiv

Diffusion Transformers (DiT) trained with flow matching in a VAE latent space have unified visual generation across images and videos. A natural next step toward a single architecture for both generation (visual synthesi…

multimodal generationText Generation

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

2026-08-05 · Junlin Han, Shengbang Tong, David Fan, Minghao Chen 외 hf

Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities int…