paper-with-me

홈 › Papers

Sentence Directed Video Object Codetection

2015-06-05 · Haonan Yu, Jeffrey Mark Siskind

We tackle the problem of video object codetection by leveraging the weak semantic constraint implied by sentences that describe the video content. Unlike most existing work that focuses on codetecting large objects which are usually salient both in size and appearance, we can codetect objects that are small or medium sized. Our method assumes no human pose or depth information such as is required by the most recent state-of-the-art method. We employ weak semantic constraint on the codetection process by pairing the video with sentences. Although the semantic information is usually simple and weak, it can greatly boost the performance of our codetection framework by reducing the search space of the hypothesized object detections. Our experiment demonstrates an average IoU score of 0.423 on a new challenging dataset which contains 15 object classes and 150 videos with 12,509 frames in total, and an average IoU score of 0.373 on a subset of an existing dataset, originally intended for activity recognition, which contains 5 object classes and 75 videos with 8,854 frames in total.

📄 PDF Abstract BibTeX arXiv:1506.02059

Code (0)

등록된 구현이 없습니다.

Tasks

Activity RecognitionObjectSentence

Similar Papers 제목 키워드 기반

Robust Object Co-detection

2013-06-01 · CVPR 2013 6 · Xin Guo, Dong Liu, Brendan Jou, Mojun Zhu 외

Object co-detection aims at simultaneous detection of objects of the same category from a pool of related images by exploiting consistent visual patterns present in candidate objects in the images. The related image set …

ClusteringObjectobject-detectionObject Detection

Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences

2020-01-19 · CVPR 2020 6 · Zhu Zhang, Zhou Zhao, Yang Zhao, Qi. Wang 외

In this paper, we consider a novel task, Spatio-Temporal Video Grounding for Multi-Form Sentences (STVG). Given an untrimmed video and a declarative/interrogative sentence depicting an object, STVG aims to localize the s…

FormObjectSentenceSpatio-Temporal Video Grounding+1

Polar Relative Positional Encoding for Video-Language Segmentation

2020-07-20 · Ke Ning, Lingxi Xie, Fei Wu, Qi Tian

In this paper, we tackle a challenging task named video-language segmentation. Given a video and a sentence in natural language, the goal is to segment the object or actor described by the sentence in video frames. To ac…

Referring Expression SegmentationSentence

Object-Aware Multi-Branch Relation Networks for Spatio-Temporal Video Grounding

2020-08-16 · Zhu Zhang, Zhou Zhao, Zhijie Lin, Baoxing Huai 외

Spatio-temporal video grounding aims to retrieve the spatio-temporal tube of a queried object according to the given sentence. Currently, most existing grounding methods are restricted to well-aligned segment-sentence pa…

DiversityObjectRelationRelation Network+3

Grounding language acquisition by training semantic parsers using captioned videos

2018-10-01 · EMNLP 2018 10 · C Ross, ace, Andrei Barbu, Yevgeni Berzak 외

We develop a semantic parser that is trained in a grounded setting using pairs of videos captioned with sentences. This setting is both data-efficient, requiring little annotation, and similar to the experience of childr…

Language AcquisitionSemantic ParsingSentence