paper-with-me

홈 › Papers

Beyond Generation: Multi-Hop Reasoning for Factual Accuracy in Vision-Language Models

2025-11-25 · Shamima Hossain arxiv

Visual Language Models (VLMs) are powerful generative tools but often produce factually inaccurate outputs due to a lack of robust reasoning capabilities. While extensive research has been conducted on integrating external knowledge for reasoning in large language models (LLMs), such efforts remain underexplored in VLMs, where the challenge is compounded by the need to bridge multiple modalities seamlessly. This work introduces a framework for knowledge-guided reasoning in VLMs, leveraging structured knowledge graphs for multi-hop verification using image-captioning task to illustrate our framework. Our approach enables systematic reasoning across multiple steps, including visual entity recognition, knowledge graph traversal, and fact-based caption refinement. We evaluate the framework using hierarchical, triple-based and bullet-point based knowledge representations, analyzing their effectiveness in factual accuracy and logical inference. Empirical results show that our approach improves factual accuracy by approximately 31% on preliminary experiments on a curated dataset of mixtures from Google Landmarks v2, Conceptual captions and Coco captions revealing key insights into reasoning patterns and failure modes. This work demonstrates the potential of integrating external knowledge for advancing reasoning in VLMs, paving the way for more reliable and knowledgable multimodal systems.

📄 PDF Abstract BibTeX arXiv:2511.20531

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge Graphs

Similar Papers 제목 키워드 기반

Beyond Factual Accuracy: Evaluating Global Reasoning Integrity in RAG Systems with LogicScore

2026-01-21 · Zhichao Yan, Yunxiao Zhao, Jiapu Wang, Jiaoyan Chen 외 arxiv

Current evaluation methods for Retrieval Augmented Generation (RAG) suffer from \textit{factual myopia}: they relentlessly emphasize factual accuracy yet neglect global logical integrity in long-form answer generation. T…

Answer Generation

Beyond Factual Correctness: Mitigating Preference-Inconsistent Explanations in Explainable Recommendation

2026-03-03 · Chengkai Wang, Baisong Liu arxiv

LLM-based explainable recommenders can produce fluent explanations that are factually correct, yet still justify items using attributes that conflict with a user's historical preferences. Such preference-inconsistent exp…

Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning

2025-07-07 · Jaedong Hwang, Kumar Tanmay, Seok-Jin Lee, Ayush Agrawal 외

Large Language Models (LLMs) have achieved strong performance in domains like mathematics, factual QA, and code generation, yet their multilingual reasoning capabilities in these tasks remain underdeveloped. Especially f…

Code Generationreinforcement-learningReinforcement Learning

Making medical vision-language models think causally across modalities with retrieval-augmented cross-modal reasoning

2026-01-26 · Weiqin Yang, Haowen Xue, Qingyi Peng, Hexuan Hu 외 arxiv

Medical vision-language models (VLMs) achieve strong performance in diagnostic reporting and image-text alignment, yet their underlying reasoning mechanisms remain fundamentally correlational, exhibiting reliance on supe…

Visual Question AnsweringMultimodal ReasoningSemantic SimilarityCausal Inference

Beyond Factual Knowledge: Benchmarking and Learning Step-Level Procedural Rule Reasoning in Large Language Models

2026-08-24 · Bohan Yu, Pengfei Cao, Chen Han, Chenxi Zhou 외 arxiv

Large language models (LLMs) excel at text understanding and generation, yet still struggle to reliably understand and apply externally provided procedural rules at scale. To evaluate this capability, we introduce RuleWo…