Natural Language Moment Retrieval
4개 벤치마크 · 논문 22편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos
Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding
LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection
Papers
DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos
Long Video Temporal Grounding (LVTG) aims at identifying specific moments within lengthy videos based on user-provided text queries for effective content retrieval. The approach taken by existing methods of dividing vide…
Natural Language Moment RetrievalNatural Language QueriesTemporal Sentence GroundingVideo GroundingLD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection
Video Moment Retrieval and Highlight Detection aim to find corresponding content in the video based on a text query. Existing models usually first use contrastive learning methods to align video and text features, then f…
Contrastive LearningDecoderHighlight DetectionMoment Retrieval+2FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding
Text-guided Video Temporal Grounding (VTG) aims to localize relevant segments in untrimmed videos based on textual descriptions, encompassing two subtasks: Moment Retrieval (MR) and Highlight Detection (HD). Although pre…
Highlight DetectionMoment RetrievalNatural Language Moment RetrievalRetrieval+1ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
Large language models (LLMs) excel at retrieving information from lengthy text, but their vision-language counterparts (VLMs) face difficulties with hour-long videos, especially for temporal grounding. Specifically, thes…
Language-Based Temporal LocalizationLanguage ModelingLanguage ModellingNatural Language Moment RetrievalLLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
Multimodal Large Language Models (MLLMs) are widely used for visual perception, understanding, and reasoning. However, long video processing and precise moment retrieval remain challenging due to LLMs' limited context si…
Moment RetrievalNatural Language Moment RetrievalRetrievalSaliency-Guided DETR for Moment Retrieval and Highlight Detection
Existing approaches for video moment retrieval and highlight detection are not able to align text and video features efficiently, resulting in unsatisfying performance and limited production usage. To address this, we pr…
Highlight DetectionMoment RetrievalNatural Language Moment RetrievalNatural Language Queries+3