paper-with-me

홈 › Papers

Improving Visual Reasoning with Iterative Evidence Refinement

2026-03-14 · Zeru Shi, Kai Mei, Yihao Quan, Dimitris N. Metaxas, Ruixiang Tang arxiv

Vision language models (VLMs) are increasingly capable of reasoning over images, but robust visual reasoning often requires re-grounding intermediate steps in the underlying visual evidence. Recent approaches typically rely on external image operations such as zooming or cropping to re-access fine-grained details during inference, which requires additional image re-encoding and can disrupt the reasoning trajectory. We argue that VLMs already provide strong internal signals for identifying and reusing visual evidence, and that these signals can be directly leveraged to support image-grounded reasoning. Motivated by this insight, we propose an end-to-end self-revisit framework, SIEVE, that trains models to re-engage image evidence through internal representations. SIEVE automatically extracts embeddings of salient image regions and injects them into the reasoning chain when additional grounding is needed, enabling later steps to condition on relevant visual cues without external tool calls or re-encoding. We use reinforcement learning to teach the model when to trigger visual revisiting and which region embeddings to retrieve and insert during the reasoning process. Experiments on multiple visual reasoning benchmarks, together with perception, reasoning, and hallucination evaluations, show that SIEVE yields consistent gains, improving performance by 8 percent on average across several benchmarks.

📄 PDF Abstract BibTeX arXiv:2603.14117

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningVisual Reasoning

Similar Papers 제목 키워드 기반

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering

2026-04-10 · Suyang Xi, Songtao Hu, Yuxiang Lai, Wangyun Dan 외 arxiv

Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely text-centric: images are encoded once as static context, and subsequent …

Visual Question AnsweringAnswer GenerationVisual Reasoning

Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning

2025-05-26 · Meng Cao, Haoze Zhao, Can Zhang, Xiaojun Chang 외

Large Vision-Language Models (LVLMs) have demonstrated impressive general capabilities across a wide range of multi-modal tasks. However, the reasoning processes of LVLMs often suffer from unreliable outputs and limited …

reinforcement-learningReinforcement LearningVisual Reasoning

PRISM: Progressive Reasoning through Iterative Slot Memory for Vision

2026-05-29 · Ziyu Wang, Shuangpeng Han, Mengmi Zhang arxiv

Modern vision models process images in a single feed-forward pass, which limits their ability to recover missing evidence or refine uncertain representations under incomplete observations. Inspired by the iterative natur…

Semantic SegmentationImage ClassificationObject Detection

FAIR-RAG: Faithful Adaptive Iterative Refinement for Retrieval-Augmented Generation

2025-10-25 · Mohammad Aghajani Asl, Majid Asgari-Bidhendi, Behrooz Minaei-Bidgoli arxiv

While Retrieval-Augmented Generation (RAG) mitigates hallucination and knowledge staleness in Large Language Models (LLMs), existing frameworks often falter on complex, multi-hop queries that require synthesizing informa…

DIVER: Dynamic Iterative Visual Evidence Reasoning for Multimodal Fake News Detection

2026-01-12 · Weilin Zhou, Zonghao Ying, Chunlei Meng, Jiahui Liu 외 arxiv

Multimodal fake news detection is crucial for mitigating adversarial misinformation. Existing methods, relying on static fusion or LLMs, face computational redundancy and hallucination risks due to weak visual foundation…

Multimodal ReasoningFake News DetectionDense Captioning