paper-with-me

Papers

EtC: Temporal Boundary Expand then Clarify for Weakly Supervised Video Grounding with Multimodal Large Language Model

2023-12-05 · Guozhang Li, Xinpeng Ding, De Cheng, Jie Li, Nannan Wang, Xinbo Gao

Early weakly supervised video grounding (WSVG) methods often struggle with incomplete boundary detection due to the absence of temporal boundary annotations. To bridge the gap between video-level and boundary-level annotation, explicit-supervision methods, i.e., generating pseudo-temporal boundaries for training, have achieved great success. However, data augmentations in these methods might disrupt critical temporal information, yielding poor pseudo boundaries. In this paper, we propose a new perspective that maintains the integrity of the original temporal content while introducing more valuable information for expanding the incomplete boundaries. To this end, we propose EtC (Expand then Clarify), first use the additional information to expand the initial incomplete pseudo boundaries, and subsequently refine these expanded ones to achieve precise boundaries. Motivated by video continuity, i.e., visual similarity across adjacent frames, we use powerful multimodal large language models (MLLMs) to annotate each frame within initial pseudo boundaries, yielding more comprehensive descriptions for expanded boundaries. To further clarify the noise of expanded boundaries, we combine mutual learning with a tailored proposal-level contrastive objective to use a learnable approach to harmonize a balance between incomplete yet clean (initial) and comprehensive yet noisy (expanded) boundaries for more precise ones. Experiments demonstrate the superiority of our method on two challenging WSVG datasets.

📄 PDF Abstract BibTeX arXiv:2312.02483

Code (0)

등록된 구현이 없습니다.

Tasks

Boundary DetectionLanguage ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelVideo Grounding

Similar Papers 제목 키워드 기반

AutoLoc: Weakly-supervised Temporal Action Localization

2018-07-22 · Zheng Shou, Hang Gao, Lei Zhang, Kazuyuki Miyazawa 외

Temporal Action Localization (TAL) in untrimmed video is important for many applications. But it is very expensive to annotate the segment-level ground truth (action class and temporal boundary). This raises the interest…

Action LocalizationTemporal Action LocalizationWeakly-supervised Temporal Action Localization

AutoLoc: Weakly-supervised Temporal Action Localization in Untrimmed Videos

2018-09-01 · ECCV 2018 9 · Zheng Shou, Hang Gao, Lei Zhang, Kazuyuki Miyazawa 외

Temporal Action Localization (TAL) in untrimmed video is important for many applications. But it is very expensive to annotate the segment-level ground truth (action class and temporal boundary). This raises the interest…

Action LocalizationTemporal Action LocalizationWeakly Supervised Action LocalizationWeakly-supervised Temporal Action Localization

Reinforcement Learning for Weakly Supervised Temporal Grounding of Natural Language in Untrimmed Videos

2020-09-18 · Jie Wu, Guanbin Li, Xiaoguang Han, Liang Lin

Temporal grounding of natural language in untrimmed videos is a fundamental yet challenging multimedia task facilitating cross-media visual content retrieval. We focus on the weakly supervised setting of this task that m…

cross-modal alignmentreinforcement-learningReinforcement Learning (RL)Retrieval+1

Mining Forgery Traces from Reconstruction Error: A Weakly Supervised Framework for Multimodal Deepfake Temporal Localization

2026-01-29 · Midou Guo, Qilin Yin, Wei Lu, Rui Yang arxiv

Modern deepfakes have evolved into localized and intermittent manipulations that require fine-grained temporal localization to mitigate severe digital security risks. The prohibitive cost of frame-level annotation makes …

Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment Localization

2026-03-16 · Cailing Han, Zhangbin Li, Jinxing Zhou, Wei Qian 외 arxiv

Point-level weakly-supervised temporal sentiment localization (P-WTSL) aims to detect sentiment-relevant segments in untrimmed multimodal videos using timestamp sentiment annotations, which greatly reduces the costly fra…

Contrastive Learning