paper-with-me

홈 › Papers

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales

2026-06-20 · Xinlong Chen, Jiafu Tang, Yue Ding, Yizhuo Jia, Bozhou Li, Bohan Zeng, Yang Shi, Shihao Li, Yiyan Ji, Qiang Liu, Weihong Lin, Yuanxing Zhang, Pengfei Wan, Liang Wang, Tieniu Tan arxiv

Accurate and comprehensive video captions with consistent subject references are critical for downstream understanding and generation tasks. However, few existing benchmarks can objectively and comprehensively evaluate these properties across diverse durations and scenarios, thereby hindering the advancement of video captioning models. To bridge this gap, we propose CapRiCorn-1K, a comprehensive benchmark designed to evaluate both video captioning quality and subject referential consistency across long temporal horizons and diverse video domains. To accommodate varied evaluation needs, our benchmark supports both audiovisual and visual-only settings. Extensive experiments on CapRiCorn-1K reveal that current models generally struggle to generate accurate and comprehensive captions while maintaining consistent subject references. Moreover, as video duration increases, both the overall caption quality and subject referential consistency decline. Notably, our evaluation metrics exhibit strong correlations with the performance of downstream understanding and generation tasks conditioned on the generated captions, further validating their effectiveness. The project is available at https://github.com/xlchen0205/CapRiCorn-1K .

📄 PDF Abstract BibTeX arXiv:2606.21949

Code (0)

등록된 구현이 없습니다.

Tasks

Video Captioning

Similar Papers 제목 키워드 기반

Spatio-Temporal Attention Models for Grounded Video Captioning

2016-10-17 · Mihai Zanfir, Elisabeta Marinoiu, Cristian Sminchisescu

Automatic video captioning is challenging due to the complex interactions in dynamic real scenes. A comprehensive system would ultimately localize and track the objects, actions and interactions present in a video and ge…

image-classificationImage ClassificationTemporal LocalizationVideo Captioning

Taking an Emotional Look at Video Paragraph Captioning

2022-03-12 · Qinyu Li, Tengpeng Li, Hanli Wang, Chang Wen Chen

Translating visual data into natural language is essential for machines to understand the world and interact with humans. In this work, a comprehensive study is conducted on video paragraph captioning, with the goal to g…

Image Captioning

SOVC: Subject-Oriented Video Captioning

2023-12-20 · Chang Teng, Yunchuan Ma, Guorong Li, Yuankai Qi 외

Describing video content according to users' needs is a long-held goal. Although existing video captioning methods have made significant progress, the generated captions may not focus on the entity that users are particu…

Video Captioning

GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration

2025-09-14 · Wan Xu, Feng Zhu, Yihan Zeng, Yuanfan Guo 외 arxiv

Video detailed captioning aims to generate comprehensive video descriptions to facilitate video understanding. Recently, most efforts in the video detailed captioning community have been made towards a local-to-global pa…

Video Captioning

CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval

2024-12-31 · Yifan Xu, Xinhao Li, Yichun Yang, Desen Meng 외

Video understanding, including video captioning and retrieval, is still a great challenge for video-language models (VLMs). The existing video retrieval and caption benchmarks only include short descriptions, limits thei…

RetrievalText RetrievalText to Video RetrievalVideo Captioning+3