Papers Natural Language Moment Retrieval
“Natural Language Moment Retrieval” 태그가 달린 논문 22편 · 필터 해제
DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos
Long Video Temporal Grounding (LVTG) aims at identifying specific moments within lengthy videos based on user-provided text queries for effective content retrieval. The approach taken by existing methods of dividing vide…
Natural Language Moment RetrievalNatural Language QueriesTemporal Sentence GroundingVideo GroundingLD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection
Video Moment Retrieval and Highlight Detection aim to find corresponding content in the video based on a text query. Existing models usually first use contrastive learning methods to align video and text features, then f…
Contrastive LearningDecoderHighlight DetectionMoment Retrieval+2FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding
Text-guided Video Temporal Grounding (VTG) aims to localize relevant segments in untrimmed videos based on textual descriptions, encompassing two subtasks: Moment Retrieval (MR) and Highlight Detection (HD). Although pre…
Highlight DetectionMoment RetrievalNatural Language Moment RetrievalRetrieval+1ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
Large language models (LLMs) excel at retrieving information from lengthy text, but their vision-language counterparts (VLMs) face difficulties with hour-long videos, especially for temporal grounding. Specifically, thes…
Language-Based Temporal LocalizationLanguage ModelingLanguage ModellingNatural Language Moment RetrievalLLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
Multimodal Large Language Models (MLLMs) are widely used for visual perception, understanding, and reasoning. However, long video processing and precise moment retrieval remain challenging due to LLMs' limited context si…
Moment RetrievalNatural Language Moment RetrievalRetrievalSaliency-Guided DETR for Moment Retrieval and Highlight Detection
Existing approaches for video moment retrieval and highlight detection are not able to align text and video features efficiently, resulting in unsatisfying performance and limited production usage. To address this, we pr…
Highlight DetectionMoment RetrievalNatural Language Moment RetrievalNatural Language Queries+3Prior Knowledge Integration via LLM Encoding and Pseudo Event Regulation for Video Moment Retrieval
In this paper, we investigate the feasibility of leveraging large language models (LLMs) for integrating general knowledge and incorporating pseudo-events as priors for temporal content distribution in video moment retri…
General KnowledgeHighlight DetectionMoment RetrievalNatural Language Moment Retrieval+2The Surprising Effectiveness of Multimodal Large Language Models for Video Moment Retrieval
Recent studies have shown promising results in utilizing multimodal large language models (MLLMs) for computer vision tasks such as object detection and semantic segmentation. However, many challenging video tasks remain…
Action LocalizationMoment RetrievalNatural Language Moment Retrievalobject-detection+4UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection
Temporal Action Detection (TAD) focuses on detecting pre-defined actions, while Moment Retrieval (MR) aims to identify the events described by open-ended natural language within untrimmed videos. Despite that they focus …
Action DetectionMoment QueriesMoment RetrievalNatural Language Moment Retrieval+2RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos
Locating specific moments within long videos (20-120 minutes) presents a significant challenge, akin to finding a needle in a haystack. Adapting existing short video (5-30 seconds) grounding methods to this problem yield…
Natural Language Moment RetrievalNatural Language QueriesRetrievalText Retrieval+1BAM-DETR: Boundary-Aligned Moment Detection Transformer for Temporal Sentence Grounding in Videos
Temporal sentence grounding aims to localize moments relevant to a language description. Recently, DETR-like approaches achieved notable progress by predicting the center and length of a target moment. However, they suff…
Moment RetrievalNatural Language Moment RetrievalSentenceTemporal Sentence GroundingBridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection
Video Moment Retrieval (MR) and Highlight Detection (HD) have attracted significant attention due to the growing demand for video analysis. Recent approaches treat MR and HD as similar video grounding problems and addres…
Contrastive LearningHighlight DetectionMoment RetrievalNatural Language Moment Retrieval+3Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding
Temporal Grounding is to identify specific moments or highlights from a video corresponding to textual descriptions. Typical approaches in temporal grounding treat all video clips equally during the encoding process rega…
Highlight DetectionMoment RetrievalNatural Language Moment RetrievalRepresentation Learning+1UnLoc: A Unified Framework for Video Localization Tasks
While large-scale image-text pretrained models such as CLIP have been used for multiple video-level tasks on trimmed videos, their use for temporal localization in untrimmed videos is still a relatively unexplored task. …
Action SegmentationMoment RetrievalNatural Language Moment RetrievalRetrieval+3UniVTG: Towards Unified Video-Language Temporal Grounding
Video Temporal Grounding (VTG), which aims to ground target clips from videos (such as consecutive intervals or disjoint shots) according to custom language queries (e.g., sentences or words), is key for video browsing o…
Highlight DetectionMoment RetrievalNatural Language Moment RetrievalRetrieval+1Background-aware Moment Detection for Video Moment Retrieval
Video moment retrieval (VMR) identifies a specific moment in an untrimmed video for a given natural language query. This task is prone to suffer the weak alignment problem innate in video datasets. Due to the ambiguity, …
Moment RetrievalNatural Language Moment RetrievalRetrievalLearning Grounded Vision-Language Representation for Versatile Understanding in Untrimmed Videos
Joint video-language learning has received increasing attention in recent years. However, existing works mainly focus on single or multiple trimmed video clips (events), which makes human-annotated event boundaries neces…
Dense Video CaptioningNatural Language Moment RetrievalSentenceText Generation+1Localizing Moments in Long Video Via Multimodal Guidance
The recent introduction of the large-scale, long-form MAD and Ego4D datasets has enabled researchers to investigate the performance of current state-of-the-art methods for video grounding in the long-form setup, with int…
Natural Language Moment RetrievalNatural Language Visual GroundingVideo GroundingVideo UnderstandingMAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio Descriptions
The recent and increasing interest in video-language research has driven the development of large-scale datasets that enable data-intensive machine learning techniques. In comparison, limited effort has been made at asse…
Moment RetrievalNatural Language Moment RetrievalVLG-Net: Video-Language Graph Matching Network for Video Grounding
Grounding language queries in videos aims at identifying the time interval (or moment) semantically relevant to a language query. The solution to this challenging task demands understanding videos' and queries' semantic …
Graph MatchingMoment RetrievalNatural Language Moment RetrievalTemporal Localization+1