OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval
Vision-language retrieval-augmented generation (RAG) has become an effective approach for tackling Knowledge-Based Visual Question Answering (KB-VQA), which requires external knowledge beyond the visual content presented in images. The effectiveness of Vision-language RAG systems hinges on multimodal retrieval, which is inherently challenging due to the diverse modalities and knowledge granularities in both queries and knowledge bases. Existing methods have not fully tapped into the potential interplay between these elements. We propose a multimodal RAG system featuring a coarse-to-fine, multi-step retrieval that harmonizes multiple granularities and modalities to enhance efficacy. Our system begins with a broad initial search aligning knowledge granularity for cross-modal retrieval, followed by a multimodal fusion reranking to capture the nuanced multimodal information for top entity selection. A text reranker then filters out the most relevant fine-grained section for augmented generation. Extensive experiments on the InfoSeek and Encyclopedic-VQA benchmarks show our method achieves state-of-the-art retrieval performance and highly competitive answering results, underscoring its effectiveness in advancing KB-VQA systems.
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-Modal RetrievalQuestion AnsweringRAGRerankingRetrievalRetrieval-augmented GenerationVisual Question AnsweringVisual Question Answering (VQA)Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Train Once, Deploy Anywhere: Matryoshka Representation Learning for Multimodal Recommendation
Despite recent advancements in language and vision modeling, integrating rich multimodal knowledge into recommender systems continues to pose significant challenges. This is primarily due to the need for efficient recomm…
Multimodal RecommendationRecommendation SystemsRepresentation LearningSequential RecommendationBloomGML: Graph Machine Learning through the Lens of Bilevel Optimization
Bilevel optimization refers to scenarios whereby the optimal solution of a lower-level energy function serves as input features to an upper-level objective of interest. These optimal features typically depend on tunable …
Bilevel OptimizationGraph LearningGraph Neural NetworkKnowledge Graph Embeddingsi-Code Studio: A Configurable and Composable Framework for Integrative AI
Artificial General Intelligence (AGI) requires comprehensive understanding and generation capabilities for a variety of tasks spanning different modalities and functionalities. Integrative AI is one important direction t…
Question AnsweringRetrievalSpeech-to-Speech TranslationText Retrieval+2SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding
Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric cues. However, the relevance of these mod…
Multimodal ReasoningScene UnderstandingMotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls
Whole-body multimodal motion generation, controlled by text, speech, or music, has numerous applications including video generation and character animation. However, employing a unified model to achieve various generatio…
Gesture GenerationMotion GenerationMotion Synthesismultimodal generation