Visual Question Answering
29개 벤치마크 · 논문 2,872편 · 이 태스크의 논문 보기 →
Benchmarks
MM-Vet
MM-Vet v2
ViP-Bench
VQA v2 test-dev
BenchLMM
MMBench
V*bench
MSRVTT-QA
VQA v2 val
VQA v2 test-std
MMHal-Bench
MSVD-QA
PlotQA-D1
PlotQA-D2
VQA v2
AID-VQA
AMBER
CLEVR
EarthVQA
GQA
GRIT
MapEval-Visual
RSVQA-HR
SIRI-WHU
TextVQA test-standard
VisualMRC
VizWiz
Most implemented
Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
VQA: Visual Question Answering
A simple neural network module for relational reasoning
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Visual Instruction Tuning
Papers
Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning
In-context learning (ICL) is widely used in multimodal large language models (MLLMs) and achieves strong performance across a wide range of multimodal tasks. However, existing multimodal ICL methods often rely on surface…
Visual Question AnsweringVDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification
Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question answering, yet they often struggle with a basic comparative skill: identifying what has changed betw…
Visual Question AnsweringImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation
We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, covering spoken visual question answering and image-grounded hallucination …
Visual Question AnsweringText-to-Image GenerationQwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D percep…
Visual Question AnsweringScene Understanding3D Object DetectionAutonomous DrivingAn Integrated Vision-and-Language Pretraining (VLP) and Visual Question Answering (VQA) model to Automate Nondestructive Evaluation Image Analysis
An AI-based approach called ChatNDE Figure to Caption is introduced, which aims to automate the interpretation of NDE images using deep learning and natural language processing (NLP). A Vision-and-Language Pretraining (V…
Visual Question AnsweringSynth-JDoc: Synthesizing a Japanese Document Image Dataset for OCR with Diverse Layouts and Embedded Images
The ability of Large Vision Language Models (LVLMs) to read text within document images is crucial, as it enables various applications such as Document Visual Question Answering. To enhance the text-reading capabilities …
Visual Question Answering