paper-with-me

Papers

RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation

2025-06-07 · Ruoxuan Zhang, Jidong Gao, Bin Wen, HongXia Xie, Chenming Zhang, Hong-Han Shuai, Wen-Huang Cheng

Creating recipe images is a key challenge in food computing, with applications in culinary education and multimodal recipe assistants. However, existing datasets lack fine-grained alignment between recipe goals, step-wise instructions, and visual content. We present RecipeGen, the first large-scale, real-world benchmark for recipe-based Text-to-Image (T2I), Image-to-Video (I2V), and Text-to-Video (T2V) generation. RecipeGen contains 26,453 recipes, 196,724 images, and 4,491 videos, covering diverse ingredients, cooking procedures, styles, and dish types. We further propose domain-specific evaluation metrics to assess ingredient fidelity and interaction modeling, benchmark representative T2I, I2V, and T2V models, and provide insights for future recipe generation models. Project page is available now.

📄 PDF Abstract BibTeX arXiv:2506.06733

Code (0)

등록된 구현이 없습니다.

Tasks

Recipe Generation

Similar Papers 제목 키워드 기반

MMDeepResearch-Bench: A Benchmark for Multimodal Deep Research Agents

2026-01-18 · Peizhou Huang, Zixuan Zhong, Zhongwei Wan, Donghao Zhou 외 arxiv

Deep Research Agents (DRAs) generate citation-rich reports via multi-step search and synthesis, yet existing benchmarks mainly target text-only settings or short-form multimodal QA, missing end-to-end multimodal evidence…

Latent-Aligned Reasoning for Multimodal Recommendation

2026-09-04 · Jiarui Jin, Anyang Ji arxiv

Multimodal Vision-Language Models (VLMs) have demonstrated remarkable capabilities in cross-modal understanding, yet a fundamental challenge persists when applying them to recommendation: as representations propagate thr…

Multimodal RecommendationContrastive Learning

EgoEMS: A High-Fidelity Multimodal Egocentric Dataset for Cognitive Assistance in Emergency Medical Services

2025-11-13 · Keshara Weerasinghe, Xueren Ge, Tessa Heick, Lahiru Nuwan Wijayasingha 외 arxiv

Emergency Medical Services (EMS) are critical to patient survival in emergencies, but first responders often face intense cognitive demands in high-stakes situations. AI cognitive assistants, acting as virtual partners, …

Speaker DiarizationDecision Making

Seed1.8 Model Card: Towards Generalized Real-World Agency

2026-03-21 · Bytedance Seed arxiv

We present Seed1.8, a foundation model aimed at generalized real-world agency: going beyond single-turn prediction to multi-turn interaction, tool use, and multi-step execution. Seed1.8 keeps strong LLM and vision-langua…

Code Generation

MC-Search: Evaluating and Enhancing Multimodal Agentic Search with Structured Long Reasoning Chains

2026-03-01 · Xuying Ning, Dongqi Fu, Tianxin Wei, Mengting Ai 외 arxiv

With the increasing demand for step-wise, cross-modal, and knowledge-grounded reasoning, multimodal large language models (MLLMs) are evolving beyond the traditional fixed retrieve-then-generate paradigm toward more soph…

Multimodal Reasoning