paper-with-me

홈 › Papers

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice

2026-07-09 · Qian Jiang, Zhecheng Shi, Jingpu Yang, Zirui Song, Miao Fang arxiv

The rapid integration of Large Vision-Language Models (VLMs) into critical infrastructure promises to revolutionize personalized healthcare and dietary management. However, in the domain of food systems, autonomous agents face a unique and persistent challenge: the "Systemic Information Asymmetry" between visual appearance and intrinsic nutritional composition. Existing benchmarks primarily focus on coarse-grained classification tasks, such as food category recognition, which fail to evaluate the intricate reasoning chain required for real-world dietary management -- specifically, the ability to traverse from identifying hidden ingredients to estimating physical mass, and finally synthesizing safety-critical medical advice. In this paper, we introduce OmniFood-Bench, a comprehensive benchmark constructed from the MM-Food-100K dataset. Unlike previous works, OmniFood-Bench evaluates VLMs across three progressive capabilities: Basic Perception (Ingredients & Cooking Methods), Quantitative Reasoning (Portion Size & Nutritional Profiling), and Safety-Critical Advisory (Disease-Specific Recommendations). We evaluate six state-of-the-art VLMs, including gpt-5.1, gemini-3-flash, and qwen3-vl-8B. Our extensive experiments reveal a startling "Semantic-Physical Gap": while models achieve near-human accuracy in naming dishes, they exhibit catastrophic failure in mass estimation and frequently hallucinate benign advice for high-risk diabetic profiles. This work establishes a rigorous standard for trustworthiness in autonomous agents deployed for public health. The code and datasets are available in: https://anonymous.4open.science/r/OmniFood-Bench-7D0B

📄 PDF Abstract BibTeX arXiv:2607.08423

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models

2025-05-13 · Pritam Sarkar, Ali Etemad

Despite recent advances in video understanding, the capabilities of Large Video Language Models (LVLMs) to perform video-based causal reasoning remains underexplored, largely due to the absence of relevant and dedicated …

FormMultiple-choiceVideo RecognitionVideo Understanding

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs

2025-05-30 · Ai Jian, Weijie Qiu, Xiaokun Wang, Peiyu Wang 외

Vision-Language Models (VLMs) have demonstrated remarkable progress in multimodal understanding, yet their capabilities for scientific reasoning remains inadequately assessed. Current multimodal benchmarks predominantly …

DiagnosticImage Comprehensionvalid

Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events

2024-12-07 · CVPR 2025 1 · Aditya Chinchure, Sahithya Ravi, Raymond Ng, Vered Shwartz 외

The commonsense reasoning capabilities of vision-language models (VLMs), especially in abductive reasoning and defeasible reasoning, remain poorly understood. Most benchmarks focus on typical visual scenarios, making it …

Beyond Accuracy: Evaluating Grounded Visual Evidence in Thinking with Images

2026-01-14 · Xuchen Li, Xuzhao Li, Renjie Pi, Shiyu Hu 외 arxiv

Despite the remarkable progress of Vision-Language Models (VLMs) in adopting "Thinking-with-Images" capabilities, accurately evaluating the authenticity of their reasoning process remains a critical challenge. Existing b…

Visual Reasoning

Oedipus and the Sphinx: Benchmarking and Improving Visual Language Models for Complex Graphic Reasoning

2025-08-01 · Jianyi Zhang, Xu Ji, Ziyin Zhou, Yuchen Zhou 외 arxiv

Evaluating the performance of visual language models (VLMs) in graphic reasoning tasks has become an important research topic. However, VLMs still show obvious deficiencies in simulating human-level graphic reasoning cap…