paper-with-me

홈 › Papers

BenchSeg: A Large-Scale Dataset and Benchmark for Multi-View Food Video Segmentation

2026-01-12 · Ahmad AlMughrabi, Guillermo Rivo, Carlos Jiménez-Farfán, Umair Haroon, Farid Al-Areqi, Hyunjun Jung, Benjamin Busam, Ricardo Marques, Petia Radeva arxiv

Food image segmentation is a critical task for dietary analysis, enabling accurate estimation of food volume and nutrients. However, current methods suffer from limited multi-view data and poor generalization to new viewpoints. We introduce BenchSeg, a novel multi-view food video segmentation dataset and benchmark. BenchSeg aggregates 55 dish scenes (from Nutrition5k, Vegetables & Fruits, MetaFood3D, and FoodKit) with 25,284 meticulously annotated frames, capturing each dish under free 360° camera motion. We evaluate a diverse set of 20 state-of-the-art segmentation models (e.g., SAM-based, transformer, CNN, and large multimodal) on the existing FoodSeg103 dataset and evaluate them (alone and combined with video-memory modules) on BenchSeg. Quantitative and qualitative results demonstrate that while standard image segmenters degrade sharply under novel viewpoints, memory-augmented methods maintain temporal consistency across frames. Our best model based on a combination of SeTR-MLA+XMem2 outperforms prior work (e.g., improving over FoodMem by ~2.63% mAP), offering new insights into food segmentation and tracking for dietary analysis. In addition to frame-wise spatial accuracy, we introduce a dedicated temporal evaluation protocol that explicitly quantifies segmentation stability over time through continuity, flicker rate, and IoU drift metrics. This allows us to reveal failure modes that remain invisible under standard per-frame evaluations. We release BenchSeg to foster future research. The project page including the dataset annotations and the food segmentation models can be found at https://amughrabi.github.io/benchseg.

📄 PDF Abstract BibTeX arXiv:2601.07581

Code (0)

등록된 구현이 없습니다.

Tasks

Image SegmentationVideo Segmentation

Similar Papers 제목 키워드 기반

PKU-MMD: A Large Scale Benchmark for Continuous Multi-Modal Human Action Understanding

2017-03-22 · Chunhui Liu, Yueyu Hu, Yanghao Li, Sijie Song 외

Despite the fact that many 3D human activity benchmarks being proposed, most existing action datasets focus on the action recognition tasks for the segmented videos. There is a lack of standard large-scale benchmarks, es…

Action DetectionAction RecognitionAction UnderstandingTemporal Action Localization

LMGQS: A Large-scale Dataset for Query-focused Summarization

2023-05-22 · Ruochen Xu, Song Wang, Yang Liu, Shuohang Wang 외

Query-focused summarization (QFS) aims to extract or generate a summary of an input document that directly answers or is relevant to a given query. The lack of large-scale datasets in the form of documents, queries, and …

DiversityLanguage ModelingLanguage ModellingQuery-focused Summarization+1

Scaling Up Deep Clustering Methods Beyond ImageNet-1K

2024-06-03 · Nikolas Adaloglou, Felix Michels, Kaspar Senft, Diana Petrusheva 외

Deep image clustering methods are typically evaluated on small-scale balanced classification datasets while feature-based $k$-means has been applied on proprietary billion-scale datasets. In this work, we explore the per…

ClusteringDeep ClusteringImage Clustering

Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training Benchmark

2022-02-14 · Jiaxi Gu, Xiaojun Meng, Guansong Lu, Lu Hou 외

Vision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datasets. However, the lack of large-scale dat…

BenchmarkingContrastive Learningimage-classificationImage Classification+6

UniKG: A Benchmark and Universal Embedding for Large-Scale Knowledge Graphs

2023-09-11 · Yide Qiu, Shaoxiang Ling, Tong Zhang, Bo Huang 외

Irregular data in real-world are usually organized as heterogeneous graphs (HGs) consisting of multiple types of nodes and edges. To explore useful knowledge from real-world data, both the large-scale encyclopedic HG dat…

AttributeGraph LearningGraph Representation LearningKnowledge Graphs+2