Highlight Detection
4개 벤치마크 · 논문 93편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries
UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding
Rhapsody: A Dataset for Highlight Detection in Podcasts
Papers
SVHighlights: Towards Extremely Long Sport Video Highlight Detection
While highlight detection for long-form videos is of great practical importance, most existing methods remain limited to short-form content, largely due to the absence of a suitable benchmark. To bridge this gap, we intr…
Highlight DetectionTuring Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
Video-language models are pivotal for tasks such as moment retrieval and highlight detection, yet they often struggle to capture the dynamic, non-linear interactions between temporal video sequences and textual semantics…
Saliency PredictionHighlight DetectionMoment RetrievalCoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection
Video Moment Retrieval (MR) and Highlight Detection (HD) are crucial tasks in video analysis that aim to localize specific moments and estimate clip-wise relevance based on a given text query. Recent approaches treat the…
Representation LearningHighlight DetectionMoment RetrievalVideo GroundingGroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
Video temporal grounding (VTG) is a critical task in video understanding and a key capability for extending video large language models (Vid-LLMs) to broader applications. However, existing Vid-LLMs rely on uniform frame…
Highlight DetectionMoment RetrievalFollow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video Captioning
Existing retrieval-augmented approaches for Dense Video Captioning (DVC) often fail to achieve accurate temporal segmentation aligned with true event boundaries, as they rely on heuristic strategies that overlook ground …
Dense Video CaptioningHighlight DetectionSounding Highlights: Dual-Pathway Audio Encoders for Audio-Visual Video Highlight Detection
Audio-visual video highlight detection aims to automatically identify the most salient moments in videos by leveraging both visual and auditory cues. However, existing models often underutilize the audio modality, focusi…
Highlight Detection