paper-with-me

홈 › Papers

GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks

2023-11-02 · Xinlu Zhang, Yujie Lu, Weizhi Wang, An Yan, Jun Yan, Lianke Qin, Heng Wang, Xifeng Yan, William Yang Wang, Linda Ruth Petzold

Automatically evaluating vision-language tasks is challenging, especially when it comes to reflecting human judgments due to limitations in accounting for fine-grained details. Although GPT-4V has shown promising results in various multi-modal tasks, leveraging GPT-4V as a generalist evaluator for these tasks has not yet been systematically explored. We comprehensively validate GPT-4V's capabilities for evaluation purposes, addressing tasks ranging from foundational image-to-text and text-to-image synthesis to high-level image-to-image translations and multi-images to text alignment. We employ two evaluation methods, single-answer grading and pairwise comparison, using GPT-4V. Notably, GPT-4V shows promising agreement with humans across various tasks and evaluation methods, demonstrating immense potential for multi-modal LLMs as evaluators. Despite limitations like restricted visual clarity grading and real-world complex reasoning, its ability to provide human-aligned scores enriched with detailed explanations is promising for universal automatic evaluator.

📄 PDF Abstract BibTeX arXiv:2311.01361

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationImage to text

Similar Papers 제목 키워드 기반

Uni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language Tasks

2022-11-17 · CVPR 2023 1 · Hao Li, Jinguo Zhu, Xiaohu Jiang, Xizhou Zhu 외

Despite the remarkable success of foundation models, their task-specific fine-tuning paradigm makes them inconsistent with the goal of general perception modeling. The key to eliminating this inconsistency is to use gene…

DecoderLanguage ModellingMulti-Task Learning

MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models

2024-07-17 · Leyang Shen, Gongwei Chen, Rui Shao, Weili Guan 외

Multimodal large language models (MLLMs) have demonstrated impressive capabilities across various vision-language tasks. However, a generalist MLLM typically underperforms compared with a specialist MLLM on most VL tasks…

MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

2023-08-04 · Weihao Yu, Zhengyuan Yang, Linjie Li, JianFeng Wang 외

We propose MM-Vet, an evaluation benchmark that examines large multimodal models (LMMs) on complicated multimodal tasks. Recent LMMs have shown various intriguing abilities, such as solving math problems written on the b…

MathMM-VetZero-Shot Visual Question Answring

Image Generators are Generalist Vision Learners

2026-04-22 · Valentin Gabeur, Shangbang Long, Songyou Peng, Paul Voigtlaender 외 arxiv

Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how LLMs develop emergent capabilities of language understanding and reasoning from generative p…

Depth EstimationImage GenerationText Generation

Vision Generalist Model: A Survey

2025-06-11 · Ziyi Wang, Yongming Rao, Shuofeng Sun, Xinrun Liu 외

Recently, we have witnessed the great success of the generalist model in natural language processing. The generalist model is a general framework trained with massive data and is able to process various downstream tasks …

modelSurvey