paper-with-me

홈 › Papers

Fine-Grained Food Image Understanding via Target-Aware Data Alignment

2026-07-28 · Jui-Feng Chi, Wei-Lun Chu, Bruce Coburn, Jinge Ma, Fengqing Zhu arxiv

Fine-grained food visual--semantic understanding requires models to capture subtle distinctions across ingredients, cooking methods, doneness, color, texture, and plate composition. Although CLIP-style vision-language models provide a natural framework for this task, their effectiveness is limited when training relies on heterogeneous web-collected image--text pairs. Such data often exhibit a web-to-target domain gap and cross-modal misalignment, where images differ from the target distribution and captions are noisy, multilingual, or weakly grounded in visual content. We propose a data-centric multimodal alignment method for fine-grained food description and recognition. Our method first performs target-aware data selection to identify visually relevant training subsets, then applies VLM-based caption refinement to generate visually grounded, target-style descriptions. Using these curated image--caption pairs, we train complementary CLIP-style retrieval experts and further combine their decisions through a hierarchical VLM-assisted multi-expert decision-level fusion strategy that invokes the VLM only when experts disagree. Experiments show that our data refinement strategy significantly improves retrieval performance over naive web supervision, with VLM-based caption refinement alone yielding an average performance gain of approximately 19%. Our full method also achieves more than twice the retrieval score of pure VLM-based retrieval while remaining substantially more efficient.

📄 PDF Abstract BibTeX arXiv:2607.25794

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GCAM: Gaussian and causal-attention model of food fine-grained recognition

2024-03-18 · Guohang Zhuang, Yue Hu, Tianxing Yan, JiaZhan Gao

Currently, most food recognition relies on deep learning for category classification. However, these approaches struggle to effectively distinguish between visually similar food samples, highlighting the pressing need to…

counterfactualCounterfactual ReasoningFine-Grained Image RecognitionFood Recognition

FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food Culture

2024-06-16 · Wenyan Li, Xinyu Zhang, Jiaang Li, Qiwei Peng 외

Food is a rich and varied dimension of cultural heritage, crucial to both individuals and social groups. To bridge the gap in the literature on the often-overlooked regional diversity in this domain, we introduce FoodieQ…

DiversityMultiple-choiceQuestion AnsweringVisual Question Answering (VQA)

A Large-Scale Benchmark for Food Image Segmentation

2021-05-12 · Xiongwei Wu, Xin Fu, Ying Liu, Ee-Peng Lim 외

Food image segmentation is a critical and indispensible task for developing health-related applications such as estimating food calories and nutrients. Existing food image segmentation models are underperforming due to t…

Image SegmentationSegmentationSemantic Segmentation

F4-ITS: Fine-grained Feature Fusion for Food Image-Text Search

2025-08-23 · Raghul Asokan arxiv

The proliferation of digital food content has intensified the need for robust and accurate systems capable of fine-grained visual understanding and retrieval. In this work, we address the challenging task of food image-t…

FoodX-251: A Dataset for Fine-grained Food Classification

2019-07-14 · Parneet Kaur, Karan Sikka, Weijun Wang, Serge Belongie 외

Food classification is a challenging problem due to the large number of categories, high visual similarity between different foods, as well as the lack of datasets for training state-of-the-art deep models. Solving this …

ClassificationFine-Grained Visual CategorizationGeneral Classification