paper-with-me

홈 › Papers

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation

2026-05-08 · Junwei Wen, Deshui Miao, Guangming Lu, Xin Li, Wenjie Pei arxiv

Video Reasoning Segmentation (VRS) aims to segment target objects in videos based on implicit instructions that convey human intent and temporal logic. Existing MLLM-based methods predict masks with a [SEG] token after selecting frames via simple sampling or an auxiliary MLLM, where limited supervision and frame-language similarity rules often yield narrow-scope keyframe choices that weaken holistic temporal understanding and lead to brittle localization in complex multi-object scenes. To address these issues, we introduce RCoT-Seg, a video-of-thought framework that factorizes VRS into temporal video reasoning (TVR) and keyframe target perception (KTP), explicitly separating temporal reasoning from spatial perception. Specifically, in the TVR stage, an agentic keyframe selection module, initialized with a curated CoT-start corpus and refined by GRPO under task-aligned rewards, is proposed to generate and reselect the keyframe through self-evaluation, strengthening moment localization and temporal reasoning. In the KTP stage, RCoT-Seg performs high-resolution segmentation on the selected frame and propagates masks with SAM2-based methods across the sequence, replacing heuristic sampling and external selectors while improving spatial precision and inter-frame consistency. Extensive experimental results demonstrate that the proposed RCoT-Seg achieves favorable performance against the state-of-the-art methods. The code and models will be publicly released at https://github.com/Victor-wjw/RCoT-Seg.

📄 PDF Abstract BibTeX arXiv:2605.07334

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RCOT: Detecting and Rectifying Factual Inconsistency in Reasoning by Reversing Chain-of-Thought

2023-05-19 · Tianci Xue, Ziqi Wang, Zhenhailong Wang, Chi Han 외

Large language Models (LLMs) have achieved promising performance on arithmetic reasoning tasks by incorporating step-by-step chain-of-thought (CoT) prompting. However, LLMs face challenges in maintaining factual consiste…

Arithmetic ReasoningGSM8KHallucination

FairCoT: Enhancing Fairness in Diffusion Models via Chain of Thought Reasoning of Multimodal Language Models

2024-06-13 · Zahraa Al Sahili, Ioannis Patras, Matthew Purver

In the domain of text-to-image generative models, biases inherent in training datasets often propagate into generated content, posing significant ethical challenges, particularly in socially sensitive contexts. We introd…

AttributeDiversityFairness

Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions

2022-12-20 · Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal

Prompting-based large language models (LLMs) are surprisingly powerful at generating natural language reasoning steps or Chains-of-Thoughts (CoT) for multi-step question answering (QA). They struggle, however, when the n…

HallucinationQuestion AnsweringRetrieval

From Generalist to Specialist: Improving Large Language Models for Medical Physics Using ARCoT

2024-05-17 · Jace Grandinetti, Rafe McBeth

Large Language Models (LLMs) have achieved remarkable progress, yet their application in specialized fields, such as medical physics, remains challenging due to the need for domain-specific knowledge. This study introduc…

BenchmarkingMultiple-choiceRetrieval

A Comprehensive Social Bias Audit of Contrastive Vision Language Models

2025-01-22 · Zahraa Al Sahili, Ioannis Patras, Matthew Purver

In the domain of text-to-image generative models, biases inherent in training datasets often propagate into generated content, posing significant ethical challenges, particularly in socially sensitive contexts. We introd…

DiversityFairnessLightweight Deployment