paper-with-me

Relational Reasoning

1개 벤치마크 · 논문 603편 · 이 태스크의 논문 보기 →

Benchmarks

CLUTRR (k=3)

결과 1개

Most implemented

Graph-Based Global Reasoning Networks

2018-11-30 · 구현 9개

Relational Deep Reinforcement Learning

2018-06-05 · 구현 7개

Papers

Do Medical Vision Models Reason About Anatomy? Probing the Spatial Inductive Biases of Learned Visual Representations

2026-08-28 · Naren Akash, Neeraja Ramanan arxiv

Interpreting a CT scan means comparing structures on either side, judging how far apart organs sit, and knowing where each one belongs. Medical vision encoders are evaluated on diagnostic accuracy, or through assembled m…

Relational Reasoning

Investigating Relational Reasoning in VLMs

2026-08-24 · Adhithya Laxman Ravi Shankar Geetha, Aulia Kharis Rakhmasari, Haleema Ramzan, Xander Yap arxiv

Vision-Language Models (VLMs) achieve strong performance in visual reasoning tasks, but it remains unclear whether they understand visual relations, or simply employ shortcuts such as language cues or priors. To investig…

Relational ReasoningVisual Reasoning

Modeling Scientific Experiment Scenes: Dataset and Model

2026-08-03 · Minghao Zou, Qingtian Zeng, Shangkun Liu, Cong Liu 외 arxiv

Scene Graph Generation (SGG) is fundamental to structured visual understanding, yet existing benchmarks focus mainly on daily life images and overlook scientific experiment scenes with specialized instruments, task-speci…

Scene Graph GenerationRelational Reasoning

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

2026-07-22 · Qiwei Ma, Chunping Qiu, Xinjun Cheng, Xiaoyu Zhang 외 arxiv

The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language interaction with remote sensing imagery…

Visual Question AnsweringRelational ReasoningScene UnderstandingVisual Grounding

Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026

2026-07-10 · Nirjhar Das, Md. Al-Mamun Provath arxiv

We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that answer pyramid-style questions…

Relational ReasoningQuestion AnsweringAnswer Selection

SGF-CDNet: A Consistency-Discrepancy Graph Network over Semantic-Geometric Fused Nodes for Face Forgery Detection

2026-07-04 · Jiayao Jiang, Bin Liu, Nenghai Yu arxiv

The rapid advancement of deepfakes necessitates robust face forgery detection. Although forged faces may lack obvious artifacts, they often contain subtle disharmony among different facial regions. We propose SGF-CDNet, …

Relational ReasoningFace Parsing

전체 603편 보기 →