paper-with-me

홈 › Papers

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding

2026-06-23 · Yixuan Li, Guangzhi Sun, Yudong Yang, Chao Zhang arxiv

Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to miss critical information for question answering (QA). A practical and efficient solution is a two-stage paradigm: first perform coarse video understanding to localize relevant segments, and then re-watch these segments at higher temporal or spatial fidelity. In this paper, we present video-SALMONN-R$^3$, the first end-to-end video-LLM that enables re-watch through reinforcement learning without relying on chain-of-thought (CoT) cold-start. This design removes the need for costly CoT data annotations and avoids CoT-based supervised fine-tuning (SFT), which can otherwise degrade the pretrained video understanding abilities. To address the mismatch between the reasoning-first behavior induced by re-watch and the answer-first tendency of pretrained video-LLMs, we propose a re-answer strategy, in which the model first produces a direct answer in the first watch and then refines it after re-watching. Finally, to improve question adherence during re-watching, we propose a re-ask mechanism that re-injects the query when revisiting localized segments. Experimental results show that video-SALMONN-R$^3$ consistently outperforms both the base model and the QA-SFT baseline, while surpassing prior re-watch-based approaches with significantly lower computational cost. Code, models, and data will be publicly released upon acceptance.

📄 PDF Abstract BibTeX arXiv:2606.24477

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningQuestion Answering

Similar Papers 제목 키워드 기반

ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis

2025-09-28 · Congzhi Zhang, Zhibin Wang, Yinchao Ma, Jiawei Peng 외 arxiv

While Reinforcement Learning with Verifiable Reward (RLVR) significantly advances image reasoning in Large Vision-Language Models (LVLMs), its application to complex video reasoning remains underdeveloped. This gap stems…

Reinforcement LearningInformation Retrieval

video-SALMONN 2: Captioning-Enhanced Audio-Visual Large Language Models

2025-06-18 · Changli Tang, Yixuan Li, Yudong Yang, Jimin Zhuang 외

Videos contain a wealth of information, and generating detailed and accurate descriptions in natural language is a key aspect of video understanding. In this paper, we present video-SALMONN 2, an advanced audio-visual la…

Audio captioningLarge Language ModelQuestion AnsweringVideo Captioning+2

video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLM

2025-10-13 · Guangzhi Sun, Yixuan Li, Xiaodong Wu, Yudong Yang 외 arxiv

Long-duration streaming video understanding is fundamental for future AI agents, yet remains limited by ineffective long-term memory. We introduce video-SALMONN S, a memory-enhanced streaming audio-visual large language …

video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

2024-06-22 · Guangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen 외

Speech understanding as an element of the more generic video understanding using audio-visual large language models (av-LLMs) is a crucial yet understudied aspect. This paper proposes video-SALMONN, a single end-to-end a…

DiversityLanguage ModelingLanguage ModellingLarge Language Model+1

Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization

2024-10-09 · Changli Tang, Yixuan Li, Yudong Yang, Jimin Zhuang 외

Videos contain a wealth of information, and generating detailed and accurate descriptions in natural language is a key aspect of video understanding. In this paper, we present video-SALMONN 2, an advanced audio-visual la…

Audio captioningLarge Language ModelQuestion AnsweringVideo Captioning+2