paper-with-me

Visual Reasoning

12개 벤치마크 · 논문 1,326편 · 이 태스크의 논문 보기 →

Benchmarks

Winoground

결과 114개

NLVR2 Dev

결과 15개

NLVR2 Test

결과 14개

CLEVRER

결과 13개

Bongard-OpenWorld

결과 9개

WinoGAViL

결과 8개

VSR

결과 5개

PHYRE-1B-Cross

결과 4개

PHYRE-1B-Within

결과 4개

VASR

결과 4개

NLVR

결과 1개

Most implemented

Visual Instruction Tuning

2023-04-17 · 구현 13개

Papers

Reason Through the Latent! Making Latent Visual Reasoning Necessary

2026-09-06 · Suhyeong Park, Junha Jung, Jaewoo Kang hf

Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than explicit textual chains of thought. However, visual information being present in a latent state does not imply tha…

Multimodal ReasoningVisual Reasoning

From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making

2026-09-04 · Davide Testa, Hugh Mee Wong, Alessandro Lenci, Bernardo Magnini 외 arxiv

Vision-Language Models are commonly evaluated through their final predictions, but understanding whether these decisions are grounded in visual evidence requires tracing how visual information contributes to language-bas…

Visual Reasoning

PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference

2026-08-27 · Junjie Liu, Shengyuan Ye, Xu Chen arxiv

Vision-Language Models (VLMs) demonstrate exceptional visual reasoning capabilities, yet their inference costs escalate rapidly with the proliferation of visual tokens. Existing visual token pruning methods exhibit two f…

Visual Reasoning

G2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification

2026-08-27 · Zehua Hao, Fang Liu, Qinliang Wang, Yaoyang Du 외 arxiv

Zero-shot classification needs efficient label retrieval and fine-grained visual reasoning, yet discriminative and generative vision-language models fail in complementary ways.When CLIP's top-1 prediction is wrong, the c…

Zero-Shot Image ClassificationVisual Reasoning

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

2026-08-26 · Junxiang Xu, Ruisi Wang, Fanyi Pu, Maijunxian Wang 외 arxiv

Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for …

Reinforcement LearningVideo GenerationVisual Reasoning

Investigating Relational Reasoning in VLMs

2026-08-24 · Adhithya Laxman Ravi Shankar Geetha, Aulia Kharis Rakhmasari, Haleema Ramzan, Xander Yap arxiv

Vision-Language Models (VLMs) achieve strong performance in visual reasoning tasks, but it remains unclear whether they understand visual relations, or simply employ shortcuts such as language cues or priors. To investig…

Relational ReasoningVisual Reasoning

전체 1,326편 보기 →