DietDelta: A Vision-Language Approach for Dietary Assessment via Before-and-After Images
Accurate dietary assessment is critical for precision nutrition, yet most image-based methods rely on a single pre-consumption image and provide only coarse, meal-level estimates. These approaches cannot determine what was actually consumed and often require restrictive inputs such as depth sensing, multi-view imagery, or explicit segmentation. In this paper, we propose a simple vision-language framework for food-item-level nutritional analysis using paired before-and-after eating images. Instead of relying on rigid segmentation masks, our method leverages natural language prompts to localize specific food items and estimate their weight directly from a single RGB image. We further estimate food consumption by predicting weight differences between paired images using a two-stage training strategy. We evaluate our method on three publicly available datasets and demonstrate consistent improvements over existing approaches, establishing a strong baseline for before-and-after dietary image analysis.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Saliency-Aware Class-Agnostic Food Image Segmentation
Advances in image-based dietary assessment methods have allowed nutrition professionals and researchers to improve the accuracy of dietary assessment, where images of food consumed are captured using smartphones or weara…
Image SegmentationNutritionSemantic SegmentationA review on vision-based analysis for automatic dietary assessment
Background: Maintaining a healthy diet is vital to avoid health-related issues, e.g., undernutrition, obesity and many non-communicable diseases. An indispensable part of the health diet is dietary assessment. Traditiona…
Food RecognitionNutritionDietary Assessment with Multimodal ChatGPT: A Systematic Analysis
Conventional approaches to dietary assessment are primarily grounded in self-reporting methods or structured interviews conducted under the supervision of dietitians. These methods, however, are often subjective, potenti…
Image CaptioningScene UnderstandingClustering Egocentric Images in Passive Dietary Monitoring with Self-Supervised Learning
In our recent dietary assessment field studies on passive dietary monitoring in Ghana, we have collected over 250k in-the-wild images. The dataset is an ongoing effort to facilitate accurate measurement of individual foo…
ClusteringSelf-Supervised LearningAn Integrated System for Mobile Image-Based Dietary Assessment
Accurate assessment of dietary intake requires improved tools to overcome limitations of current methods including user burden and measurement error. Emerging technologies such as image-based approaches using advanced ma…
Nutrition