Visual Question Answering (VQA)
76개 벤치마크 · 논문 2,167편 · 이 태스크의 논문 보기 →
Benchmarks
GQA Test2019
VQA v2 test-dev
VQA v2 test-std
OK-VQA
MSVD-QA
MSRVTT-QA
DocVQA test
InfographicVQA
GQA test-dev
VizWiz 2020 VQA
A-OKVQA
CLEVR
InfiMM-Eval
IconQA
TextVQA test-standard
VCR (Q-A) test
VQA v2 val
VQA-CP
VizWiz 2018
VLM2-Bench
VQA-CE
VCR (QA-R) test
GQA test-std
IllusionVQA
InfoSeek
VCR (Q-AR) test
VQA v1 test-dev
VQA v1 test-std
WHOOPS!
AutoHallusion
CLEVR-Humans
QLEVR
AI2D
HallusionBench
PMC-VQA
PlotQA-D1
PlotQA-D2
Visual7W
F-VQA
FigureQA - test 1
VCR (Q-A) dev
VCR (Q-AR) dev
VCR (QA-R) dev
DocVQA
DocVQA val
GQA
GRIT
TDIUC
TGIF-QA
VQA-X
ActivityNet
ArtQuest
COCO
CORE-MM
DVQA test-familiar
DeepForm
EgoSchema
ImageNet
MM-Vet
MME
MVBench
OVAD benchmark
RetVQA
TextVQA
Video MME
Visual Genome (pairs)
Visual Genome (subjects)
WebSRC
ZS-F-VQA
Most implemented
Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
ParlAI: A Dialog Research Software Platform
VQA: Visual Question Answering
A simple neural network module for relational reasoning
Papers
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
Recent advancements in vision-language models (VLMs) have improved performance by increasing the number of visual tokens, which are often significantly longer than text tokens. However, we observe that most real-world sc…
Language ModelingLanguage ModellingOptical Character Recognition (OCR)reinforcement-learning+2MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM
Recent studies have utilized visual large language models (VLMs) to answer not only "Is this face a forgery?" but also "Why is the face a forgery?" These studies introduced forgery-related attributes, such as forgery loc…
AttributeFace SwappingPrompt LearningVisual Question Answering (VQA)Describe Anything Model for Visual Question Answering on Text-rich Images
Recent progress has been made in region-aware vision-language modeling, particularly with the emergence of the Describe Anything Model (DAM). DAM is capable of generating detailed descriptions of any specific image areas…
DescriptiveLanguage ModelingLanguage ModellingQuestion Answering+2Evaluating Attribute Confusion in Fashion Text-to-Image Generation
Despite the rapid advances in Text-to-Image (T2I) generation models, their evaluation remains challenging in domains like fashion, involving complex compositional generation. Recent automated T2I evaluation methods lever…
Attributecross-modal alignmentImage GenerationQuestion Answering+5LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation
Large Multimodal Models (LMMs) are typically trained on vast corpora of image-text data but are often limited in linguistic coverage, leading to biased and unfair outputs across languages. While prior work has explored m…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder
Existing video segmenter and grounder approaches, exemplified by Sa2VA, directly fuse features within segmentation models. This often results in an undesirable entanglement of dynamic visual information and static semant…
Image SegmentationLarge Language ModelQuestion AnsweringSegmentation+5