Summarizing First-Person Videos from Third Persons' Points of Views
Video highlight or summarization is among interesting topics in computer vision, which benefits a variety of applications like viewing, searching, or storage. However, most existing studies rely on training data of third-person videos, which cannot easily generalize to highlight the first-person ones. With the goal of deriving an effective model to summarize first-person videos, we propose a novel deep neural network architecture for describing and discriminating vital spatiotemporal information across videos with different points of view. Our proposed model is realized in a semi-supervised setting, in which fully annotated third-person videos, unlabeled first-person videos, and a small number of annotated first-person ones are presented during training. In our experiments, qualitative and quantitative evaluations on both benchmarks and our collected first-person video datasets are presented.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Summarizing First-Person Videos from Third Persons' Points of View
Video highlight or summarization is among interesting topics in computer vision, which benefits a variety of applications like viewing, searching, or storage. However, most existing studies rely on training data of third…
Third person enforcement in a prisoner's dilemma game
We theoretically study the effect of a third person enforcement on a one-shot prisoner's dilemma game played by two persons, with whom the third person plays repeated prisoner's dilemma games. We find that the possibilit…
Actor and Observer: Joint Modeling of First and Third-Person Videos
Several theories in cognitive neuroscience suggest that when people interact with the world, or simulate interactions, they do so from a first-person egocentric perspective, and seamlessly transfer knowledge between thir…
Action RecognitionTemporal Action LocalizationIdentifying First-person Camera Wearers in Third-person Videos
We consider scenarios in which we wish to perform joint scene understanding, object tracking, activity recognition, and other tasks in environments in which multiple people are wearing body-worn cameras while a third-per…
Activity RecognitionObject TrackingScene UnderstandingTripletVictory Sign Biometric for Terrorists Identification
Covering the face and all body parts, sometimes the only evidence to identify a person is their hand geometry, and not the whole hand- only two fingers (the index and the middle fingers) while showing the victory sign, a…