paper-with-me

홈 › Papers

SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information

2024-09-21 · Jiashuo Sun, Jihai Zhang, Yucheng Zhou, Zhaochen Su, Xiaoye Qu, Yu Cheng

Large Vision-Language Models (LVLMs) have become pivotal at the intersection of computer vision and natural language processing. However, the full potential of LVLMs Retrieval-Augmented Generation (RAG) capabilities remains underutilized. Existing works either focus solely on the text modality or are limited to specific tasks. Moreover, most LVLMs struggle to selectively utilize retrieved information and are sensitive to irrelevant or misleading references. To address these challenges, we propose a self-refinement framework designed to teach LVLMs to Selectively Utilize Retrieved Information (SURf). Specifically, when given questions that are incorrectly answered by the LVLM backbone, we obtain references that help correct the answers (positive references) and those that do not (negative references). We then fine-tune the LVLM backbone using a combination of these positive and negative references. Our experiments across three tasks and seven datasets demonstrate that our framework significantly enhances LVLMs ability to effectively utilize retrieved multimodal references and improves their robustness against irrelevant or misleading information. The source code is available at https://github.com/GasolSun36/SURf.

📄 PDF Abstract BibTeX arXiv:2409.14083

Code (1)

gasolsun36/surf 공식 구현 pytorch

Tasks

RAGRetrieval-augmented Generation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

ORCESTRA: VLM-driven Visual Robot programming in Mixed Reality

2026-08-01 · Ivan Snegirev, Elizaveta Semenyakina, Mikhail Konenkov, Artem Lykov 외 arxiv

ORCESTRA is a mixed-reality system for programming robot digital twins through no-code waypoint teaching and language-guided control. In a passthrough mixed-reality workspace, users place robot twins on real surfaces, te…

TeachArena: Are Language Agents Ready for Realistic Teaching Work?

2026-05-14 · Zixin Chen, Peng Liu, Rui Sheng, Haobo Li 외 arxiv

Language agents are increasingly deployed in professional workflows, yet tutoring remains a high-stakes capability that existing evaluations only partially capture. Effective tutor agents require more than producing corr…

UCO: A Multi-Turn Interactive Reinforcement Learning Method for Adaptive Teaching with Large Language Models

2025-11-12 · Shouang Wei, Min Zhang, Xin Lin, Bo Jiang 외 arxiv

Large language models (LLMs) are shifting from answer providers to intelligent tutors in educational settings, yet current supervised fine-tuning methods only learn surface teaching patterns without dynamic adaptation ca…

Reinforcement Learning

Teaching Structured Vision&Language Concepts to Vision&Language Models

2022-11-21 · Sivan Doveh, Assaf Arbelle, Sivan Harary, Rameswar Panda 외

Vision and Language (VL) models have demonstrated remarkable zero-shot performance in a variety of tasks. However, some aspects of complex language understanding still remain a challenge. We introduce the collective noti…

Teaching Structured Vision & Language Concepts to Vision & Language Models

2023-01-01 · CVPR 2023 1 · Sivan Doveh, Assaf Arbelle, Sivan Harary, Eli Schwartz 외

Vision and Language (VL) models have demonstrated remarkable zero-shot performance in a variety of tasks. However, some aspects of complex language understanding still remain a challenge. We introduce the collective …