paper-with-me

Papers

OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval

2025-05-10 · Wei Yang, Jingjing Fu, Rui Wang, Jinyu Wang, Lei Song, Jiang Bian

Vision-language retrieval-augmented generation (RAG) has become an effective approach for tackling Knowledge-Based Visual Question Answering (KB-VQA), which requires external knowledge beyond the visual content presented in images. The effectiveness of Vision-language RAG systems hinges on multimodal retrieval, which is inherently challenging due to the diverse modalities and knowledge granularities in both queries and knowledge bases. Existing methods have not fully tapped into the potential interplay between these elements. We propose a multimodal RAG system featuring a coarse-to-fine, multi-step retrieval that harmonizes multiple granularities and modalities to enhance efficacy. Our system begins with a broad initial search aligning knowledge granularity for cross-modal retrieval, followed by a multimodal fusion reranking to capture the nuanced multimodal information for top entity selection. A text reranker then filters out the most relevant fine-grained section for augmented generation. Extensive experiments on the InfoSeek and Encyclopedic-VQA benchmarks show our method achieves state-of-the-art retrieval performance and highly competitive answering results, underscoring its effectiveness in advancing KB-VQA systems.

📄 PDF Abstract BibTeX arXiv:2505.07879

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalQuestion AnsweringRAGRerankingRetrievalRetrieval-augmented GenerationVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Train Once, Deploy Anywhere: Matryoshka Representation Learning for Multimodal Recommendation

2024-09-25 · Yueqi Wang, Zhenrui Yue, Huimin Zeng, Dong Wang 외

Despite recent advancements in language and vision modeling, integrating rich multimodal knowledge into recommender systems continues to pose significant challenges. This is primarily due to the need for efficient recomm…

Multimodal RecommendationRecommendation SystemsRepresentation LearningSequential Recommendation

BloomGML: Graph Machine Learning through the Lens of Bilevel Optimization

2024-03-07 · Amber Yijia Zheng, Tong He, Yixuan Qiu, Minjie Wang 외

Bilevel optimization refers to scenarios whereby the optimal solution of a lower-level energy function serves as input features to an upper-level objective of interest. These optimal features typically depend on tunable …

Bilevel OptimizationGraph LearningGraph Neural NetworkKnowledge Graph Embeddings

i-Code Studio: A Configurable and Composable Framework for Integrative AI

2023-05-23 · Yuwei Fang, Mahmoud Khademi, Chenguang Zhu, ZiYi Yang 외

Artificial General Intelligence (AGI) requires comprehensive understanding and generation capabilities for a variety of tasks spanning different modalities and functionalities. Integrative AI is one important direction t…

Question AnsweringRetrievalSpeech-to-Speech TranslationText Retrieval+2

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

2026-08-05 · Yue Zhang, Yingzhao Jian, Yunqiu Xu, Xiaoxiao Sun 외 hf

Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric cues. However, the relevance of these mod…

Multimodal ReasoningScene Understanding

MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls

2024-07-30 · Yuxuan Bian, Ailing Zeng, Xuan Ju, Xian Liu 외

Whole-body multimodal motion generation, controlled by text, speech, or music, has numerous applications including video generation and character animation. However, employing a unified model to achieve various generatio…

Gesture GenerationMotion GenerationMotion Synthesismultimodal generation