paper-with-me

홈 › Papers

MV-CoRe: Multimodal Visual-Conceptual Reasoning for Complex Visual Question Answering

2025-08-09 · Jingwei Peng, Jiehao Chen, Mateo Alejandro Rojas, Meilin Zhang arxiv

Complex Visual Question Answering (Complex VQA) tasks, which demand sophisticated multi-modal reasoning and external knowledge integration, present significant challenges for existing large vision-language models (LVLMs) often limited by their reliance on high-level global features. To address this, we propose MV-CoRe (Multimodal Visual-Conceptual Reasoning), a novel model designed to enhance Complex VQA performance through the deep fusion of diverse visual and linguistic information. MV-CoRe meticulously integrates global embeddings from pre-trained Vision Large Models (VLMs) and Language Large Models (LLMs) with fine-grained semantic-aware visual features, including object detection characteristics and scene graph representations. An innovative Multimodal Fusion Transformer then processes and deeply integrates these diverse feature sets, enabling rich cross-modal attention and facilitating complex reasoning. We evaluate MV-CoRe on challenging Complex VQA benchmarks, including GQA, A-OKVQA, and OKVQA, after training on VQAv2. Our experimental results demonstrate that MV-CoRe consistently outperforms established LVLM baselines, achieving an overall accuracy of 77.5% on GQA. Ablation studies confirm the critical contribution of both object and scene graph features, and human evaluations further validate MV-CoRe's superior factual correctness and reasoning depth, underscoring its robust capabilities for deep visual and conceptual understanding.

📄 PDF Abstract BibTeX arXiv:2508.07023

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringObject Detection

Similar Papers 제목 키워드 기반

Visual Abstract Thinking Empowers Multimodal Reasoning

2025-05-26 · Dairu Liu, Ziyue Wang, Minyuan Ruan, Fuwen Luo 외

Images usually convey richer detail than text, but often include redundant information which potentially downgrades multimodal reasoning performance. When faced with lengthy or complex messages, humans tend to employ abs…

Multimodal ReasoningRelational ReasoningVisual Reasoning

M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding

2026-08-06 · Hong Jiang, Junnan Zhu, Jingwang Huang, Xiao Sun 외 arxiv

Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information jointly construct Target--Source mappings…

Reinforcement Learning

DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models

2025-12-30 · Zefeng He, Xiaoye Qu, Yafu Li, Tong Zhu 외 arxiv

While recent Multimodal Large Language Models (MLLMs) have attained significant strides in multimodal reasoning, their reasoning processes remain predominantly text-centric, leading to suboptimal performance in complex l…

Multimodal Reasoning

Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning

2025-11-30 · Haozhen Gong, Xiaozhong Ji, Yuansen Liu, Wenbin Wu 외 arxiv

MLLMs MLLMs are beginning to appear in clinical workflows, but their ability to perform complex medical reasoning remains unclear. We present Med-CMR, a fine-grained Medical Complex Multimodal Reasoning benchmark. Med-CM…

Visual Question AnsweringMultimodal ReasoningObject Detection

Large Language Models Achieve Gold Medal Performance at the International Olympiad on Astronomy & Astrophysics (IOAA)

2025-10-06 · Lucas Carrit Delgado Pinheiro, Ziru Chen, Bruno Caixeta Piazza, Ness Shroff 외 arxiv

While task-specific demonstrations show early success in applying large language models (LLMs) to automate some astronomical research tasks, they only provide incomplete views of all necessary capabilities in solving ast…