paper-with-me

Dense Video Captioning

4개 벤치마크 · 논문 90편 · 이 태스크의 논문 보기 →

Benchmarks

ActivityNet Captions

결과 24개

YouCook2

결과 14개

ViTT

결과 8개

VidChapters-7M

결과 2개

Most implemented

Multi-modal Dense Video Captioning

2020-03-17 · 구현 4개

SoccerNet 2023 Challenges Results

2023-09-12 · 구현 2개

Papers

Claim-Level Rubric Rewards for Video Caption Reinforcement Learning

2026-07-06 · Mingqi Gao, Hongyuan Dong, Yifei Chen, Zhisheng Zhong 외 arxiv

In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinforcement learning for dense video captioning. Existing reward designs …

Dense Video CaptioningReinforcement Learning

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning

2026-07-03 · Wenzheng Zeng, Siyi Jiao, Chen Gao, Hwee Tou Ng 외 hf

Dense video captioning aims to generate temporally grounded descriptions of video events, benefiting both event-level video understanding and generation. In this domain, autoregressive video large language models have em…

Dense Video Captioning

CodecCap: High-Fidelity Codec-Inspired Residual Modeling for Dense Video Captioning

2026-05-26 · Zihan Lin, Songhe Deng, Shuwei He, Danxiang Zhu 외 arxiv

Existing video captioning methods struggle to balance visual fidelity and redundancy: holistic captions are compact but lose fine-grained evidence, whereas segment-wise captions improve coverage but introduce heavy redun…

Dense Video CaptioningDense Captioning

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation

2026-04-29 · Mingji Ge, Qirui Chen, Zeqian Li, Weidi Xie arxiv

Long-term video understanding requires interpreting complex temporal events and reasoning over procedural activities. While instructional video corpora, like HowTo100M, offer rich resources for model training, they prese…

Zero-shot GeneralizationDense Video CaptioningCross-Modal Retrieval

SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusion

2026-03-13 · Xiang Li, Heqian Qiu, Lanxiao Wang, Benliu Qiu 외 arxiv

Error detection is crucial in industrial training, healthcare, and assembly quality control. Most existing work assumes a single-view setting and cannot handle the practical case where a third-person (exo) demonstration …

Dense Video CaptioningAction Detection

Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video Captioning

2026-03-12 · Seung Hyup Baek, Jimin Lee, Hyeongkeun Lee, Jae Won Cho arxiv

Dense Video Captioning (DVC) is a challenging multimodal task that involves temporally localizing multiple events within a video and describing them with natural language. While query-based frameworks enable the simultan…

Dense Video Captioning

전체 90편 보기 →