paper-with-me

Dense Captioning

1개 벤치마크 · 논문 86편 · 이 태스크의 논문 보기 →

Benchmarks

Visual Genome

결과 4개

Most implemented

Dense-Captioning Events in Videos

2017-05-02 · 구현 4개

Papers

VisChronos: Revolutionizing Image Captioning Through Real-Life Events

2026-06-23 · Phuc-Tan Nguyen, Hieu Nguyen, Minh-Triet Tran, Trung-Nghia Le arxiv

This paper aims to bridge the semantic gap between visual content and natural language understanding by leveraging historical events in the real world as a source of knowledge for caption generation. We propose VisChrono…

Natural Language UnderstandingDense CaptioningImage Captioning

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning

2026-06-08 · Penghui Yang, Long Xing, Xiaoyi Dong, Yuhang Zang 외 arxiv

Image and video captioning are fundamental tasks that bridge the visual and linguistic domains, playing a critical role in pre-training Large Vision-Language Models (LVLMs). Current state-of-the-art captioning models are…

Reinforcement LearningVideo CaptioningDense Captioning

CodecCap: High-Fidelity Codec-Inspired Residual Modeling for Dense Video Captioning

2026-05-26 · Zihan Lin, Songhe Deng, Shuwei He, Danxiang Zhu 외 arxiv

Existing video captioning methods struggle to balance visual fidelity and redundancy: holistic captions are compact but lose fine-grained evidence, whereas segment-wise captions improve coverage but introduce heavy redun…

Dense Video CaptioningDense Captioning

COPRA: Conditional Parameter Adaptation with Reinforcement Learning for Video Anomaly Detection

2026-05-14 · Darryl Cherian Jacob, Xinyu Liu, Kai Wang, Pan He arxiv

Vision-language models (VLMs) have shown strong performance in video anomaly detection (VAD) while providing interpretable predictions. However, existing VLM-based VAD methods suffer from a fundamental mismatch between t…

Video Question AnsweringVideo Anomaly DetectionReinforcement LearningDense Captioning

Curvature-Aware Captioning:Leveraging Geodesic Attention for 3D Scene Understanding

2026-05-09 · Ziyao He, Yingjie Liu, ZhangYangRui, Mingsong Chen 외 arxiv

Accurate 3D scene description is fundamental to robotic navigation and augmented reality, yet current dense captioning methods face significant limitations in processing sparse point cloud data. % Existing approaches tha…

Object LocalizationScene UnderstandingDense Captioning

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding

2026-04-28 · Minghang Zheng, Zihao Yin, Yi Yang, Yuxin Peng 외 arxiv

Video Temporal Grounding (VTG), the task of localizing video segments from text queries, struggles in open-world settings due to limited dataset scale and semantic diversity, causing performance gaps between common and r…

Reinforcement LearningDense Captioning

전체 86편 보기 →