paper-with-me

Papers

FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food Culture

2024-06-16 · Wenyan Li, Xinyu Zhang, Jiaang Li, Qiwei Peng, Raphael Tang, Li Zhou, Weijia Zhang, Guimin Hu, Yifei Yuan, Anders Søgaard, Daniel Hershcovich, Desmond Elliott

Food is a rich and varied dimension of cultural heritage, crucial to both individuals and social groups. To bridge the gap in the literature on the often-overlooked regional diversity in this domain, we introduce FoodieQA, a manually curated, fine-grained image-text dataset capturing the intricate features of food cultures across various regions in China. We evaluate vision-language Models (VLMs) and large language models (LLMs) on newly collected, unseen food images and corresponding questions. FoodieQA comprises three multiple-choice question-answering tasks where models need to answer questions based on multiple images, a single image, and text-only descriptions, respectively. While LLMs excel at text-based question answering, surpassing human accuracy, the open-sourced VLMs still fall short by 41% on multi-image and 21% on single-image VQA tasks, although closed-weights models perform closer to human levels (within 10%). Our findings highlight that understanding food and its cultural implications remains a challenging and under-explored direction.

📄 PDF Abstract BibTeX arXiv:2406.11030

Code (1)

lyan62/FoodieQA 공식 구현 pytorch

Tasks

DiversityMultiple-choiceQuestion AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

FG-CLIP: Fine-Grained Visual and Textual Alignment

2025-05-08 · Chunyu Xie, Bin Wang, Fanjing Kong, Jincheng Li 외

Contrastive Language-Image Pre-training (CLIP) excels in multimodal tasks such as image-text retrieval and zero-shot classification but struggles with fine-grained understanding due to its focus on coarse-grained short c…

Image-text Retrievalobject-detectionObject DetectionOpen-vocabulary object detection+5

LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

2025-01-14 · CVPR 2025 1 · Hongyu Li, Jinyu Chen, Ziyu Wei, Shaofei Huang 외

Recent advancements in multimodal large language models (MLLMs) have shown promising results, yet existing approaches struggle to effectively handle both temporal and spatial localization simultaneously. This challenge s…

Feature CompressionLanguage ModelingLanguage ModellingLarge Language Model+3

Enhancing Conceptual Understanding in Multimodal Contrastive Learning through Hard Negative Samples

2024-03-05 · Philipp J. Rösch, Norbert Oswald, Michaela Geierhos, Jindřich Libovický

Current multimodal models leveraging contrastive learning often face limitations in developing fine-grained conceptual understanding. This is due to random negative samples during pretraining, causing almost exclusively …

Concept AlignmentContrastive LearningImage-text Retrieval

ERNIE-mmLayout: Multi-grained MultiModal Transformer for Document Understanding

2022-09-18 · Wenjin Wang, Zhengjie Huang, Bin Luo, Qianglong Chen 외

Recent efforts of multimodal Transformers have improved Visually Rich Document Understanding (VrDU) tasks via incorporating visual and textual information. However, existing approaches mainly focus on fine-grained elemen…

Common Sense Reasoningdocument understandingQuestion Answering

Autonomous Workflow for Multimodal Fine-Grained Training Assistants Towards Mixed Reality

2024-05-16 · Jiahuan Pei, Irene Viola, Haochen Huang, Junxiao Wang 외

Autonomous artificial intelligence (AI) agents have emerged as promising protocols for automatically understanding the language-based environment, particularly with the exponential development of large language models (L…

Mixed RealityQuestion Answering