paper-with-me

Visual Question Answering

29개 벤치마크 · 논문 2,872편 · 이 태스크의 논문 보기 →

Benchmarks

MM-Vet

결과 462개

MM-Vet v2

결과 48개

ViP-Bench

결과 26개

VQA v2 test-dev

결과 22개

BenchLMM

결과 20개

MMBench

결과 10개

V*bench

결과 10개

MSRVTT-QA

결과 8개

VQA v2 val

결과 8개

VQA v2 test-std

결과 6개

MMHal-Bench

결과 4개

MSVD-QA

결과 4개

PlotQA-D1

결과 4개

PlotQA-D2

결과 4개

VQA v2

결과 4개

AID-VQA

결과 2개

AMBER

결과 2개

CLEVR

결과 2개

EarthVQA

결과 2개

GQA

결과 2개

GRIT

결과 2개

MapEval-Visual

결과 2개

RSVQA-HR

결과 2개

SIRI-WHU

결과 2개

TextVQA test-standard

결과 2개

VisualMRC

결과 2개

VizWiz

결과 2개

Most implemented

VQA: Visual Question Answering

2015-05-03 · 구현 21개

Visual Instruction Tuning

2023-04-17 · 구현 13개

Papers

Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning

2026-09-09 · Mingbo Yang, Wenqiang Wang, Zhaolu Kang, Peng Chen 외 arxiv

In-context learning (ICL) is widely used in multimodal large language models (MLLMs) and achieves strong performance across a wide range of multimodal tasks. However, existing multimodal ICL methods often rely on surface…

Visual Question Answering

VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification

2026-09-05 · Yixin Wan, Tianle Zheng, Kai-Wei Chang hf

Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question answering, yet they often struggle with a basic comparative skill: identifying what has changed betw…

Visual Question Answering

ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

2026-08-31 · Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti, Farina Amir 외 arxiv

We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, covering spoken visual question answering and image-grounded hallucination …

Visual Question AnsweringText-to-Image Generation

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

2026-08-31 · Xin Zhou, Zongchuang Zhao, Zhibo Yang, Mingsheng Li 외 hf

We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D percep…

Visual Question AnsweringScene Understanding3D Object DetectionAutonomous Driving

An Integrated Vision-and-Language Pretraining (VLP) and Visual Question Answering (VQA) model to Automate Nondestructive Evaluation Image Analysis

2026-08-29 · Mehrdad Shafiei Dizaji, Hoda Azari arxiv

An AI-based approach called ChatNDE Figure to Caption is introduced, which aims to automate the interpretation of NDE images using deep learning and natural language processing (NLP). A Vision-and-Language Pretraining (V…

Visual Question Answering

Synth-JDoc: Synthesizing a Japanese Document Image Dataset for OCR with Diverse Layouts and Embedded Images

2026-08-28 · Keito Sasagawa, Shuhei Kurita, Daisuke Kawahara arxiv

The ability of Large Vision Language Models (LVLMs) to read text within document images is crucial, as it enables various applications such as Document Visual Question Answering. To enhance the text-reading capabilities …

Visual Question Answering

전체 2,872편 보기 →