paper-with-me

Papers

EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization

2026-07-06 · Youngkil Song, Yoonjae Baek, Dongwon Kim, Inho Kim, Dongkeun Kim, Suha Kwak arxiv

Reasoning temporal localization (RTL) requires a model to generate an answer that itself contains the time interval supporting it, so high-level reasoning and precise temporal grounding must be produced jointly in a single response. To tackle this challenging task, we propose the first event-centric video chain-of-thought framework, dubbed EventCoT. EventCoT first performs event-centric tokenization of the input video to convert it into compact event tokens, enabling efficient identification of question-relevant events. It then reasons within the identified events to generate the answer, grounding the time interval via embedding matching that aligns placeholder tokens with visual embeddings. EventCoT achieves state-of-the-art results on ActivityNet-RTL for reasoning temporal localization while using substantially fewer visual tokens than previous work. To verify its general performance, we further evaluate EventCoT on the grounded video question answering benchmark ReXTime, where it attains strong zero-shot results.

📄 PDF Abstract BibTeX arXiv:2607.04872

Code (0)

등록된 구현이 없습니다.

Tasks

Video Question Answering

Similar Papers 제목 키워드 기반

Chain of Event-Centric Causal Thought for Physically Plausible Video Generation

2026-03-10 · Zixuan Wang, Yixin Hu, Haolan Wang, Feng Chen 외 arxiv

Physically Plausible Video Generation (PPVG) has emerged as a promising avenue for modeling real-world physical phenomena. PPVG requires an understanding of commonsense knowledge, which remains a challenge for video diff…

Video Generation

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning

2025-06-16 · Shulin Tian, Ruiqi Wang, Hongming Guo, Penghao Wu 외

We introduce Ego-R1, a novel framework for reasoning over ultra-long (i.e., in days and weeks) egocentric videos, which leverages a structured Chain-of-Tool-Thought (CoTT) process, orchestrated by an Ego-R1 Agent trained…

Reinforcement Learning (RL)

Open Set Video HOI detection from Action-Centric Chain-of-Look Prompting

2023-01-01 · ICCV 2023 1 · Nan Xi, Jingjing Meng, Junsong Yuan

Human-Object Interaction (HOI) detection is essential for understanding and modeling real-world events. Existing works on HOI detection mainly focus on static images and a closed setting, where all HOI classes are pr…

Human-Object Interaction DetectionLanguage ModellingVisual Reasoning

HiVid-Narrator: Hierarchical Video Narrative Generation with Scene-Primed ASR-anchored Compression

2026-01-12 · Haoxuan Li, Mengyan Li, Junjun Zheng arxiv

Generating structured narrations for real-world e-commerce videos requires models to perceive fine-grained visual details and organize them into coherent, high-level stories--capabilities that existing approaches struggl…

Video Captioning

EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT

2025-10-27 · Baoqi Pei, Yifei Huang, Jilan Xu, Yuping He 외 arxiv

Egocentric video reasoning centers on an unobservable agent behind the camera who dynamically shapes the environment, requiring inference of hidden intentions and recognition of fine-grained interactions. This core chall…