paper-with-me

홈 › Papers

Understanding ME? Multimodal Evaluation for Fine-grained Visual Commonsense

2022-11-10 · Zhecan Wang, Haoxuan You, Yicheng He, Wenhao Li, Kai-Wei Chang, Shih-Fu Chang

Visual commonsense understanding requires Vision Language (VL) models to not only understand image and text but also cross-reference in-between to fully integrate and achieve comprehension of the visual scene described. Recently, various approaches have been developed and have achieved high performance on visual commonsense benchmarks. However, it is unclear whether the models really understand the visual scene and underlying commonsense knowledge due to limited evaluation data resources. To provide an in-depth analysis, we present a Multimodal Evaluation (ME) pipeline to automatically generate question-answer pairs to test models' understanding of the visual scene, text, and related knowledge. We then take a step further to show that training with the ME data boosts the model's performance in standard VCR evaluation. Lastly, our in-depth analysis and comparison reveal interesting findings: (1) semantically low-level information can assist the learning of high-level information but not the opposite; (2) visual information is generally under utilization compared with text.

📄 PDF Abstract BibTeX arXiv:2211.05895

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding

2024-02-28 · Yihao Ding, Lorenzo Vaiani, Caren Han, Jean Lee 외

This paper presents a groundbreaking multimodal, multi-task, multi-teacher joint-grained knowledge distillation model for visually-rich form document understanding. The model is designed to leverage insights from both fi…

document understandingFormKnowledge Distillation

MARINER: A 3E-Driven Benchmark for Fine-Grained Perception and Complex Reasoning in Open-Water Environments

2026-04-09 · Xingming Liao, Ning Chen, Muying Shu, Yunpeng Yin 외 arxiv

Fine-grained visual understanding and high-level reasoning in real-world open-water environments remain under-explored due to the lack of dedicated benchmarks. We introduce MARINER, a comprehensive benchmark built under …

Visual Question AnsweringObject Detection

LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants

2025-07-07 · Haochen Huang, Jiahuan Pei, Mohammad Aliannejadi, Xin Sun 외 arxiv

Vision-language models (VLMs) are facing the challenges of understanding and following multimodal assembly instructions, particularly when fine-grained spatial reasoning and precise object state detection are required. I…

Spatial ReasoningObject Detection

ERNIE-mmLayout: Multi-grained MultiModal Transformer for Document Understanding

2022-09-18 · Wenjin Wang, Zhengjie Huang, Bin Luo, Qianglong Chen 외

Recent efforts of multimodal Transformers have improved Visually Rich Document Understanding (VrDU) tasks via incorporating visual and textual information. However, existing approaches mainly focus on fine-grained elemen…

Common Sense Reasoningdocument understandingQuestion Answering

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling

2026-08-09 · Jinbo Yan, Limeng Qiao, Jie Qin, Junyan He 외 hf

Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation. However, their final tokens discard fine-grained visual details, leading to poor…

Text-to-Image GenerationImage ReconstructionImage Editing