paper-with-me

홈 › Papers

Food-R1: A Unified Multi-Task Food Vision-Language Model with Reinforcement Learning

2026-06-03 · Yu Zhu, Yongkang Li, Wenjie Zhu, Haoyi Jiang, Wenyu Liu, Wei Yang, Bin Li, Xinggang Wang arxiv

Recent studies have explored Vision-Language Models (VLMs) for food analysis. However, most existing methods rely primarily on supervised fine-tuning (SFT), which often limits reasoning and generalization capabilities. Moreover, high-quality large-scale nutritional annotations remain scarce. To address these issues, we introduce CalorieBench-80K, a large-scale benchmark with curated calorie labels and dietary advice annotations. To the best of our knowledge, it is the first food image benchmark to incorporate Chain-of-Thought (CoT) annotations for calorie reasoning. We also propose Food-R1, a unified food VLM trained in a multi-task learning paradigm to equip the model with broad capabilities. Food-R1 undergoes CoT-based cold-start instruction tuning, followed by reinforcement fine-tuning (RFT) using Group Relative Policy Optimization (GRPO) to improve reasoning and performance. Experiments on CalorieBench-80K and representative benchmarks show that Food-R1 consistently outperforms strong baselines across food-related tasks. The code, model weights, and benchmark annotations are available at the project repository.

📄 PDF Abstract BibTeX arXiv:2606.04986

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMulti-Task Learning

Similar Papers 제목 키워드 기반

RoDE: Linear Rectified Mixture of Diverse Experts for Food Large Multi-Modal Models

2024-07-17 · Pengkun Jiao, Xinlan Wu, Bin Zhu, Jingjing Chen 외

Large Multi-modal Models (LMMs) have significantly advanced a variety of vision-language tasks. The scalability and availability of high-quality training data play a pivotal role in the success of LMMs. In the realm of f…

GPUNutrition

Large Scale Visual Food Recognition

2021-03-30 · Weiqing Min, Zhiling Wang, Yuxin Liu, Mengjiang Luo 외

Food recognition plays an important role in food choice and intake, which is essential to the health and well-being of humans. It is thus of importance to the computer vision community, and can further support many food-…

Fine-Grained Visual RecognitionFood RecognitionImage RetrievalRepresentation Learning+1

Improving Food Image Recognition with Noisy Vision Transformer

2025-03-24 · Tonmoy Ghosh, Edward Sazonov

Food image recognition is a challenging task in computer vision due to the high variability and complexity of food images. In this study, we investigate the potential of Noisy Vision Transformers (NoisyViT) for improving…

Food Recognition

FoodLMM: A Versatile Food Assistant using Large Multi-modal Model

2023-12-22 · Yuehao Yin, Huiyan Qi, Bin Zhu, Jingjing Chen 외

Large Multi-modal Models (LMMs) have made impressive progress in many vision-language tasks. Nevertheless, the performance of general LMMs in specific domains is still far from satisfactory. This paper proposes FoodLMM, …

Food RecognitionMulti-Task LearningNutritionReasoning Segmentation+2

Applications of knowledge graphs for food science and industry

2021-07-13 · Weiqing Min, Chunlin Liu, Leyi Xu, Shuqiang Jiang

The deployment of various networks (e.g., Internet of Things [IoT] and mobile networks), databases (e.g., nutrition tables and food compositional databases), and social media (e.g., Instagram and Twitter) generates huge …

Data Visualizationgraph constructionKnowledge GraphsNutrition+2