paper-with-me

Natural Language Moment Retrieval

4개 벤치마크 · 논문 22편 · 이 태스크의 논문 보기 →

Benchmarks

TACoS

결과 13개

ActivityNet Captions

결과 8개

MAD

결과 8개

DiDeMo

결과 1개

Most implemented

Papers

DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos

2025-05-22 · CVPR 2025 1 · Zijia Lu, A S M Iftekhar, Gaurav Mittal, Tianjian Meng 외

Long Video Temporal Grounding (LVTG) aims at identifying specific moments within lengthy videos based on user-provided text queries for effective content retrieval. The approach taken by existing methods of dividing vide…

Natural Language Moment RetrievalNatural Language QueriesTemporal Sentence GroundingVideo Grounding

LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection

2025-01-18 · Pengcheng Zhao, Zhixian He, Fuwei Zhang, Shujin Lin 외

Video Moment Retrieval and Highlight Detection aim to find corresponding content in the video based on a text query. Existing models usually first use contrastive learning methods to align video and text features, then f…

Contrastive LearningDecoderHighlight DetectionMoment Retrieval+2

FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding

2024-12-18 · Zhuo Cao, Bingqing Zhang, Heming Du, Xin Yu 외

Text-guided Video Temporal Grounding (VTG) aims to localize relevant segments in untrimmed videos based on textual descriptions, encompassing two subtasks: Moment Retrieval (MR) and Highlight Detection (HD). Although pre…

Highlight DetectionMoment RetrievalNatural Language Moment RetrievalRetrieval+1

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos

2024-11-22 · CVPR 2025 1 · Tanveer Hannan, Md Mohaiminul Islam, Jindong Gu, Thomas Seidl 외

Large language models (LLMs) excel at retrieving information from lengthy text, but their vision-language counterparts (VLMs) face difficulties with hour-long videos, especially for temporal grounding. Specifically, thes…

Language-Based Temporal LocalizationLanguage ModelingLanguage ModellingNatural Language Moment Retrieval

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

2024-11-21 · Weiheng Lu, Jian Li, An Yu, Ming-Ching Chang 외

Multimodal Large Language Models (MLLMs) are widely used for visual perception, understanding, and reasoning. However, long video processing and precise moment retrieval remain challenging due to LLMs' limited context si…

Moment RetrievalNatural Language Moment RetrievalRetrieval

Saliency-Guided DETR for Moment Retrieval and Highlight Detection

2024-10-02 · Aleksandr Gordeev, Vladimir Dokholyan, Irina Tolstykh, Maksim Kuprashevich

Existing approaches for video moment retrieval and highlight detection are not able to align text and video features efficiently, resulting in unsatisfying performance and limited production usage. To address this, we pr…

Highlight DetectionMoment RetrievalNatural Language Moment RetrievalNatural Language Queries+3

전체 22편 보기 →