paper-with-me

Papers

Weak Supervision and Referring Attention for Temporal-Textual Association Learning

2020-06-21 · Zhiyuan Fang, Shu Kong, Zhe Wang, Charless Fowlkes, Yezhou Yang

A system capturing the association between video frames and textual queries offer great potential for better video analysis. However, training such a system in a fully supervised way inevitably demands a meticulously curated video dataset with temporal-textual annotations. Therefore we provide a Weak-Supervised alternative with our proposed Referring Attention mechanism to learn temporal-textual association (dubbed WSRA). The weak supervision is simply a textual expression (e.g., short phrases or sentences) at video level, indicating this video contains relevant frames. The referring attention is our designed mechanism acting as a scoring function for grounding the given queries over frames temporally. It consists of multiple novel losses and sampling strategies for better training. The principle in our designed mechanism is to fully exploit 1) the weak supervision by considering informative and discriminative cues from intra-video segments anchored with the textual query, 2) multiple queries compared to the single video, and 3) cross-video visual similarities. We validate our WSRA through extensive experiments for temporally grounding by languages, demonstrating that it outperforms the state-of-the-art weakly-supervised methods notably.

📄 PDF Abstract BibTeX arXiv:2006.11747

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ViewSAM: Learning View-aware Cross-modal Semantics for Weakly Supervised Cross-view Referring Multi-Object Tracking

2026-05-04 · Jiawei Ge, Xintian Zhang, Jiuxin Cao, Bo Liu 외 arxiv

Cross-view Referring Multi-Object Tracking (CRMOT) aims to track multiple objects specified by natural language across multiple camera views, with globally consistent identities. Despite recent progress, existing methods…

Multi-Object Tracking

Knowledge-guided Pairwise Reconstruction Network for Weakly Supervised Referring Expression Grounding

2019-09-05 · Xuejing Liu, Liang Li, Shuhui Wang, Zheng-Jun Zha 외

Weakly supervised referring expression grounding (REG) aims at localizing the referential entity in an image according to linguistic query, where the mapping between the image region (proposal) and the query is unknown i…

ObjectReferring ExpressionRegion ProposalWeakly-supervised Learning

GroPrompt: Efficient Grounded Prompting and Adaptation for Referring Video Object Segmentation

2024-06-18 · Ci-Siang Lin, I-Jieh Liu, Min-Hung Chen, Chien-Yi Wang 외

Referring Video Object Segmentation (RVOS) aims to segment the object referred to by the query sentence throughout the entire video. Most existing methods require end-to-end training with dense mask annotations, which co…

Contrastive LearningObjectPositionPrompt Learning+6

Weakly-Supervised Referring Video Object Segmentation through Text Supervision

2026-04-20 · Miaojing Shi, Jun Huang, Zijie Yue, Hanli Wang arxiv

Referring video object segmentation (RVOS) aims to segment the target instance in a video, referred by a text expression. Conventional approaches are mostly supervised learning, requiring expensive pixel-level mask annot…

Referring Video Object SegmentationReferring Expression

Learning Structured Visual Compositional Representations for Weakly Supervised Referring Expression Comprehension

2026-07-06 · Lian Xu, Mohammed Bennamoun, Farid Boussaid, Hamid Laga 외 arxiv

Referring expression comprehension (REC) aims to localize the object in an image described by natural language. In Weakly supervised REC (WREC), existing approaches primarily operate on anchor-level visual representation…

Referring Expression