paper-with-me

Visual Question Answering (VQA)

76개 벤치마크 · 논문 2,167편 · 이 태스크의 논문 보기 →

Benchmarks

GQA Test2019

결과 127개

VQA v2 test-dev

결과 56개

VQA v2 test-std

결과 38개

OK-VQA

결과 37개

MSVD-QA

결과 36개

MSRVTT-QA

결과 34개

DocVQA test

결과 33개

InfographicVQA

결과 21개

GQA test-dev

결과 17개

VizWiz 2020 VQA

결과 16개

A-OKVQA

결과 15개

CLEVR

결과 15개

InfiMM-Eval

결과 14개

IconQA

결과 12개

TextVQA test-standard

결과 12개

VCR (Q-A) test

결과 11개

VQA v2 val

결과 11개

VQA-CP

결과 10개

VizWiz 2018

결과 10개

VLM2-Bench

결과 9개

VQA-CE

결과 9개

VCR (QA-R) test

결과 8개

GQA test-std

결과 7개

IllusionVQA

결과 7개

InfoSeek

결과 7개

VCR (Q-AR) test

결과 7개

VQA v1 test-dev

결과 7개

VQA v1 test-std

결과 6개

WHOOPS!

결과 6개

AutoHallusion

결과 5개

CLEVR-Humans

결과 5개

QLEVR

결과 5개

AI2D

결과 4개

HallusionBench

결과 4개

PMC-VQA

결과 4개

PlotQA-D1

결과 4개

PlotQA-D2

결과 4개

Visual7W

결과 4개

F-VQA

결과 3개

FigureQA - test 1

결과 3개

VCR (Q-A) dev

결과 3개

VCR (Q-AR) dev

결과 3개

VCR (QA-R) dev

결과 3개

DocVQA

결과 2개

DocVQA val

결과 2개

GQA

결과 2개

GRIT

결과 2개

TDIUC

결과 2개

TGIF-QA

결과 2개

VQA-X

결과 2개

ActivityNet

결과 1개

ArtQuest

결과 1개

COCO

결과 1개

CORE-MM

결과 1개

DVQA test-familiar

결과 1개

DeepForm

결과 1개

EgoSchema

결과 1개

ImageNet

결과 1개

MM-Vet

결과 1개

MME

결과 1개

MVBench

결과 1개

OVAD benchmark

결과 1개

RetVQA

결과 1개

TextVQA

결과 1개

Video MME

결과 1개

Visual Genome (pairs)

결과 1개

WebSRC

결과 1개

ZS-F-VQA

결과 1개

Most implemented

ParlAI: A Dialog Research Software Platform

2017-05-18 · 구현 23개

VQA: Visual Question Answering

2015-05-03 · 구현 21개

Papers

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning

2025-07-17 · Senqiao Yang, Junyi Li, Xin Lai, Bei Yu 외

Recent advancements in vision-language models (VLMs) have improved performance by increasing the number of visual tokens, which are often significantly longer than text tokens. However, we observe that most real-world sc…

Language ModelingLanguage ModellingOptical Character Recognition (OCR)reinforcement-learning+2

MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM

2025-07-16 · Tao Chen, Jingyi Zhang, Decheng Liu, Chunlei Peng

Recent studies have utilized visual large language models (VLMs) to answer not only "Is this face a forgery?" but also "Why is the face a forgery?" These studies introduced forgery-related attributes, such as forgery loc…

AttributeFace SwappingPrompt LearningVisual Question Answering (VQA)

Describe Anything Model for Visual Question Answering on Text-rich Images

2025-07-16 · Yen-Linh Vu, Dinh-Thang Duong, Truong-Binh Duong, Anh-Khoi Nguyen 외

Recent progress has been made in region-aware vision-language modeling, particularly with the emergence of the Describe Anything Model (DAM). DAM is capable of generating detailed descriptions of any specific image areas…

DescriptiveLanguage ModelingLanguage ModellingQuestion Answering+2

Evaluating Attribute Confusion in Fashion Text-to-Image Generation

2025-07-09 · Ziyue Liu, Federico Girella, Yiming Wang, Davide Talon

Despite the rapid advances in Text-to-Image (T2I) generation models, their evaluation remains challenging in domains like fashion, involving complex compositional generation. Recent automated T2I evaluation methods lever…

Attributecross-modal alignmentImage GenerationQuestion Answering+5

LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation

2025-07-09 · Ananya Raval, Aravind Narayanan, Vahid Reza Khazaie, Shaina Raza

Large Multimodal Models (LMMs) are typically trained on vast corpora of image-text data but are often limited in linguistic coverage, leading to biased and unfair outputs across languages. While prior work has explored m…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder

2025-06-28 · Dang Jisheng, Wu Xudong, Wang Bimei, Lv Ning 외

Existing video segmenter and grounder approaches, exemplified by Sa2VA, directly fuse features within segmentation models. This often results in an undesirable entanglement of dynamic visual information and static semant…

Image SegmentationLarge Language ModelQuestion AnsweringSegmentation+5

전체 2,167편 보기 →