paper-with-me

홈 › Papers

Video ChatCaptioner: Towards Enriched Spatiotemporal Descriptions

2023-04-09 · Jun Chen, Deyao Zhu, Kilichbek Haydarov, Xiang Li, Mohamed Elhoseiny

Video captioning aims to convey dynamic scenes from videos using natural language, facilitating the understanding of spatiotemporal information within our environment. Although there have been recent advances, generating detailed and enriched video descriptions continues to be a substantial challenge. In this work, we introduce Video ChatCaptioner, an innovative approach for creating more comprehensive spatiotemporal video descriptions. Our method employs a ChatGPT model as a controller, specifically designed to select frames for posing video content-driven questions. Subsequently, a robust algorithm is utilized to answer these visual queries. This question-answer framework effectively uncovers intricate video details and shows promise as a method for enhancing video content. Following multiple conversational rounds, ChatGPT can summarize enriched video content based on previous conversations. We qualitatively demonstrate that our Video ChatCaptioner can generate captions containing more visual details about the videos. The code is publicly available at https://github.com/Vision-CAIR/ChatCaptioner

📄 PDF Abstract BibTeX arXiv:2304.04227

Code (1)

vision-cair/chatcaptioner 공식 구현 pytorch

Tasks

Video Captioning

Similar Papers 제목 키워드 기반

ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

2023-03-12 · Deyao Zhu, Jun Chen, Kilichbek Haydarov, Xiaoqian Shen 외

Asking insightful questions is crucial for acquiring knowledge and expanding our understanding of the world. However, the importance of questioning has been largely overlooked in AI research, where models have been prima…

Image CaptioningQuestion AnsweringVisual Reasoning

Context-Enriched Natural Language Descriptions of Vessel Trajectories

2026-03-08 · Kostas Patroumpas, Alexandros Troupiotis-Kapeliaris, Giannis Spiliopoulos, Panagiotis Betchavas 외 arxiv

We address the problem of transforming raw vessel trajectory data collected from AIS into structured and semantically enriched representations interpretable by humans and directly usable by machine reasoning systems. We …

iMOVE: Instance-Motion-Aware Video Understanding

2025-02-17 · Jiaze Li, Yaya Shi, Zongyang Ma, Haoran Xu 외

Enhancing the fine-grained instance spatiotemporal motion perception capabilities of Video Large Language Models is crucial for improving their temporal and general video understanding. However, current models struggle t…

Computational EfficiencyVideo Understanding

SeMob: Semantic Synthesis for Dynamic Urban Mobility Prediction

2025-09-24 · Runfei Chen, Shuyang Jiang, Wei Huang arxiv

Human mobility prediction is vital for urban services, but often fails to account for abrupt changes from external events. Existing spatiotemporal models struggle to leverage textual descriptions detailing these events. …

Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting

2025-04-07 · Yunlong Tang, Jing Bi, Chao Huang, Susan Liang 외

We present CAT-V (Caption AnyThing in Video), a training-free framework for fine-grained object-centric video captioning that enables detailed descriptions of user-selected objects through time. CAT-V integrates three ke…

Boundary DetectionObjectSemantic SegmentationVideo Captioning