Dense Captioning
1개 벤치마크 · 논문 86편 · 이 태스크의 논문 보기 →
Benchmarks
Visual Genome
Most implemented
3D-LLM: Injecting the 3D World into Large Language Models
Dense-Captioning Events in Videos
A Hierarchical Approach for Generating Descriptive Image Paragraphs
ComiCap: A VLMs pipeline for dense captioning of Comic Panels
Papers
VisChronos: Revolutionizing Image Captioning Through Real-Life Events
This paper aims to bridge the semantic gap between visual content and natural language understanding by leveraging historical events in the real world as a source of knowledge for caption generation. We propose VisChrono…
Natural Language UnderstandingDense CaptioningImage CaptioningCapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning
Image and video captioning are fundamental tasks that bridge the visual and linguistic domains, playing a critical role in pre-training Large Vision-Language Models (LVLMs). Current state-of-the-art captioning models are…
Reinforcement LearningVideo CaptioningDense CaptioningCodecCap: High-Fidelity Codec-Inspired Residual Modeling for Dense Video Captioning
Existing video captioning methods struggle to balance visual fidelity and redundancy: holistic captions are compact but lose fine-grained evidence, whereas segment-wise captions improve coverage but introduce heavy redun…
Dense Video CaptioningDense CaptioningCOPRA: Conditional Parameter Adaptation with Reinforcement Learning for Video Anomaly Detection
Vision-language models (VLMs) have shown strong performance in video anomaly detection (VAD) while providing interpretable predictions. However, existing VLM-based VAD methods suffer from a fundamental mismatch between t…
Video Question AnsweringVideo Anomaly DetectionReinforcement LearningDense CaptioningCurvature-Aware Captioning:Leveraging Geodesic Attention for 3D Scene Understanding
Accurate 3D scene description is fundamental to robotic navigation and augmented reality, yet current dense captioning methods face significant limitations in processing sparse point cloud data. % Existing approaches tha…
Object LocalizationScene UnderstandingDense CaptioningOmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
Video Temporal Grounding (VTG), the task of localizing video segments from text queries, struggles in open-world settings due to limited dataset scale and semantic diversity, causing performance gaps between common and r…
Reinforcement LearningDense Captioning