paper-with-me

Papers

Weakly Supervised Food Image Segmentation using Vision Transformers and Segment Anything Model

2025-09-23 · Ioannis Sarafis, Alexandros Papadopoulos, Anastasios Delopoulos arxiv

In this paper, we propose a weakly supervised semantic segmentation approach for food images which takes advantage of the zero-shot capabilities and promptability of the Segment Anything Model (SAM) along with the attention mechanisms of Vision Transformers (ViTs). Specifically, we use class activation maps (CAMs) from ViTs to generate prompts for SAM, resulting in masks suitable for food image segmentation. The ViT model, a Swin Transformer, is trained exclusively using image-level annotations, eliminating the need for pixel-level annotations during training. Additionally, to enhance the quality of the SAM-generated masks, we examine the use of image preprocessing techniques in combination with single-mask and multi-mask SAM generation strategies. The methodology is evaluated on the FoodSeg103 dataset, generating an average of 2.4 masks per image (excluding background), and achieving an mIoU of 0.54 for the multi-mask scenario. We envision the proposed approach as a tool to accelerate food image annotation tasks or as an integrated component in food and nutrition tracking applications.

📄 PDF Abstract BibTeX arXiv:2509.19028

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SegmentationImage Segmentation

Similar Papers 제목 키워드 기반

Food Image Classification and Segmentation with Attention-based Multiple Instance Learning

2023-08-22 · Valasia Vlachopoulou, Ioannis Sarafis, Alexandros Papadopoulos

The demand for accurate food quantification has increased in the recent years, driven by the needs of applications in dietary monitoring. At the same time, computer vision approaches have exhibited great potential in aut…

image-classificationImage ClassificationMultiple Instance LearningSemantic Segmentation

Combining Weakly and Webly Supervised Learning for Classifying Food Images

2017-12-23 · Parneet Kaur, Karan Sikka, Ajay Divakaran

Food classification from images is a fine-grained classification problem. Manual curation of food images is cost, time and scalability prohibitive. On the other hand, web data is available freely but contains noise. In t…

ClassificationGeneral ClassificationWeakly-supervised Learning

A Comprehensive Analysis of Weakly-Supervised Semantic Segmentation in Different Image Domains

2019-12-24 · Lyndon Chan, Mahdi S. Hosseini, Konstantinos N. Plataniotis

Recently proposed methods for weakly-supervised semantic segmentation have achieved impressive performance in predicting pixel classes despite being trained with only image labels which lack positional information. Becau…

SegmentationSemantic SegmentationWeakly supervised Semantic SegmentationWeakly-Supervised Semantic Segmentation

Probabilistic Graphlet Cut: Exploiting Spatial Structure Cue for Weakly Supervised Image Segmentation

2013-06-01 · CVPR 2013 6 · Luming Zhang, Mingli Song, Zicheng Liu, Xiao Liu 외

Weakly supervised image segmentation is a challenging problem in computer vision field. In this paper, we present a new weakly supervised image segmentation algorithm by learning the distribution of spatially structured …

Image SegmentationSegmentationSemantic SegmentationSuperpixels

NoPeopleAllowed: The Three-Step Approach to Weakly Supervised Semantic Segmentation

2020-06-13 · Mariia Dobko, Ostap Viniavskyi, Oles Dobosevych

We propose a novel approach to weakly supervised semantic segmentation, which consists of three consecutive steps. The first two steps extract high-quality pseudo masks from image-level annotated data, which are then use…

Missing LabelsSegmentationSemantic SegmentationWeakly supervised Semantic Segmentation+1