paper-with-me

Papers

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum

2026-03-05 · Shan Ning, Longtian Qiu, Xuming He arxiv

Knowledge-Based Visual Question Answering (KB-VQA) requires models to answer questions about an image by integrating external knowledge, posing significant challenges due to noisy retrieval and the structured, encyclopedic nature of the knowledge base. These characteristics create a distributional gap from pretrained multimodal large language models (MLLMs), making effective reasoning and domain adaptation difficult in the post-training stage. In this work, we propose \textit{Wiki-R1}, a data-generation-based curriculum reinforcement learning framework that systematically incentivizes reasoning in MLLMs for KB-VQA. Wiki-R1 constructs a sequence of training distributions aligned with the model's evolving capability, bridging the gap from pretraining to the KB-VQA target distribution. We introduce \textit{controllable curriculum data generation}, which manipulates the retriever to produce samples at desired difficulty levels, and a \textit{curriculum sampling strategy} that selects informative samples likely to yield non-zero advantages during RL updates. Sample difficulty is estimated using observed rewards and propagated to unobserved samples to guide learning. Experiments on two KB-VQA benchmarks, Encyclopedic VQA and InfoSeek, demonstrate that Wiki-R1 achieves new state-of-the-art results, improving accuracy from 35.5\% to 37.1\% on Encyclopedic VQA and from 40.1\% to 44.1\% on InfoSeek. The project page is available at https://artanic30.github.io/project_pages/WikiR1/.

📄 PDF Abstract BibTeX arXiv:2603.05256

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringReinforcement LearningMultimodal ReasoningDomain Adaptation

Similar Papers 제목 키워드 기반

ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval

2025-07-29 · Nicola Fanelli, Gennaro Vessio, Giovanna Castellano arxiv

Analyzing digitized artworks presents unique challenges, requiring not only visual interpretation but also a deep understanding of rich artistic, contextual, and historical knowledge. We introduce ArtSeek, a multimodal f…

Visual Question AnsweringMultimodal Reasoning

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

2026-08-03 · Huanyao Zhang, Jiepeng Zhou, Runhao Zhao, Yanzhe Shan 외 hf

Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world pro…

Reinforcement Learning

VL-KGE: Vision-Language Models Meet Knowledge Graph Embeddings

2026-03-02 · Athanasios Efthymiou, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg 외 arxiv

Real-world multimodal knowledge graphs (MKGs) are inherently heterogeneous, modeling entities that are associated with diverse modalities. Traditional knowledge graph embedding (KGE) methods excel at learning continuous …

Knowledge Graph EmbeddingKnowledge GraphsLink Prediction

Proactive Reasoning-with-Retrieval Framework for Medical Multimodal Large Language Models

2025-10-21 · Lehan Wang, Yi Qin, Honglong Yang, Xiaomeng Li arxiv

Incentivizing the reasoning ability of Multimodal Large Language Models (MLLMs) is essential for medical applications to transparently analyze medical scans and provide reliable diagnosis. However, existing medical MLLMs…

Reinforcement Learning

DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry

2025-12-12 · Zhenyang Cai, Jiaming Zhang, Junjie Zhao, Ziyi Zeng 외 arxiv

Reliable interpretation of multimodal data in dentistry is essential for automated oral healthcare, yet current multimodal large language models (MLLMs) struggle to capture fine-grained dental visual details and lack suf…

Reinforcement Learning