paper-with-me

홈 › Papers

Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs

2024-04-23 · Davide Caffagni, Federico Cocchi, Nicholas Moratelli, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

Multimodal LLMs are the natural evolution of LLMs, and enlarge their capabilities so as to work beyond the pure textual modality. As research is being carried out to design novel architectures and vision-and-language adapters, in this paper we concentrate on endowing such models with the capability of answering questions that require external knowledge. Our approach, termed Wiki-LLaVA, aims at integrating an external knowledge source of multimodal documents, which is accessed through a hierarchical retrieval pipeline. Relevant passages, using this approach, are retrieved from the external knowledge source and employed as additional context for the LLM, augmenting the effectiveness and precision of generated dialogues. We conduct extensive experiments on datasets tailored for visual question answering with external data and demonstrate the appropriateness of our approach.

📄 PDF Abstract BibTeX arXiv:2404.15406

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringRetrievalRetrieval-augmented GenerationVisual Question Answering

Similar Papers 제목 키워드 기반

Hierarchical Retrieval-Augmented Generation Model with Rethink for Multi-hop Question Answering

2024-08-20 · XiaoMing Zhang, Ming Wang, Xiaocui Yang, Daling Wang 외

Multi-hop Question Answering (QA) necessitates complex reasoning by integrating multiple pieces of information to resolve intricate questions. However, existing QA systems encounter challenges such as outdated informatio…

Multi-hop Question AnsweringQuestion AnsweringRetrievalRetrieval-augmented Generation

HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation

2025-11-19 · Linyin Luo, Yujuan Ding, Yunshan Ma, Wenqi Fan 외 arxiv

Advanced multimodal Retrieval-Augmented Generation (MRAG) techniques have been widely applied to enhance the capabilities of Large Multimodal Models (LMMs), but they also bring along novel safety issues. Existing adversa…

LLaVA Needs More Knowledge: Retrieval Augmented Natural Language Generation with Knowledge Graph for Explaining Thoracic Pathologies

2024-10-07 · Ameer Hamza, Abdullah, Yong Hyun Ahn, Sungyoung Lee 외

Generating Natural Language Explanations (NLEs) for model predictions on medical images, particularly those depicting thoracic pathologies, remains a critical and challenging task. Existing methodologies often struggle d…

RAGRetrievalText Generation

REVERSUM: A Multi-staged Retrieval-Augmented Generation Method to Enhance Wikipedia Tail Biographies through Personal Narratives

2025-02-17 · Sayantan Adak, Pauras Mangesh Meher, Paramita Das, Animesh Mukherjee

Wikipedia is an invaluable resource for factual information about a wide range of entities. However, the quality of articles on less-known entities often lags behind that of the well-known ones. This study proposes a nov…

ArticlesInformativenessRetrieval-augmented Generation

TrajWiki: Source-Grounded Memory Trajectories for Long-Horizon Dialogue Agents

2026-08-02 · Jingyu Sun, Yuyang Xue, Mingyang Li, Zhengtao Yao 외 arxiv

Large language model agents have shown strong capabilities in generating coherent and contextually appropriate responses, yet robust long-horizon dialogue remains limited by the lack of external memory that is traceable,…

Answer Generation