paper-with-me

Papers

Learning Features by Watching Objects Move

2016-12-19 · CVPR 2017 7 · Deepak Pathak, Ross Girshick, Piotr Dollár, Trevor Darrell, Bharath Hariharan

This paper presents a novel yet intuitive approach to unsupervised feature learning. Inspired by the human visual system, we explore whether low-level motion-based grouping cues can be used to learn an effective visual representation. Specifically, we use unsupervised motion-based segmentation on videos to obtain segments, which we use as 'pseudo ground truth' to train a convolutional network to segment objects from a single frame. Given the extensive evidence that motion plays a key role in the development of the human visual system, we hope that this straightforward approach to unsupervised learning will be more effective than cleverly designed 'pretext' tasks studied in the literature. Indeed, our extensive experiments show that this is the case. When used for transfer learning on object detection, our representation significantly outperforms previous unsupervised approaches across multiple settings, especially when training data for the target task is scarce.

📄 PDF Abstract BibTeX arXiv:1612.06370

Code (1)

pathak22/unsupervised-video 공식 구현 torch

Tasks

object-detectionObject DetectionTransfer Learning

Similar Papers 제목 키워드 기반

SmartBullets: A Cloud-Assisted Bullet Screen Filter based on Deep Learning

2019-05-15 · Haoran Niu, Jiangnan Li, Yu Zhao

Bullet-screen is a technique that enables the website users to send real-time comment `bullet' cross the screen. Compared with the traditional review of a video, bullet-screen provides new features of feeling expression …

Robots with Different Embodiments Can Express and Influence Carefulness in Object Manipulation

2022-08-03 · Linda Lastrico, Luca Garello, Francesco Rea, Nicoletta Noceti 외

Humans have an extraordinary ability to communicate and read the properties of objects by simply watching them being carried by someone else. This level of communicative skills and interpretation, available to humans, is…

Object

DOVE: Learning Deformable 3D Objects by Watching Videos

2021-07-22 · Shangzhe Wu, Tomas Jakab, Christian Rupprecht, Andrea Vedaldi

Learning deformable 3D objects from 2D images is often an ill-posed problem. Existing methods rely on explicit supervision to establish multi-view correspondences, such as template shape models and keypoint annotations, …

Human-to-Robot Interaction: Learning from Video Demonstration for Robot Imitation

2026-02-22 · Thanh Nguyen Canh, Thanh-Tuan Tran, Haolan Zhang, Ziyan Gao 외 arxiv

Learning from Demonstration (LfD) offers a promising paradigm for robot skill acquisition. Recent approaches attempt to extract manipulation commands directly from video demonstrations, yet face two critical challenges: …

Reinforcement LearningAction ClassificationRobot ManipulationVideo Captioning

Edge Based Grid Super-Imposition for Crowd Emotion Recognition

2016-08-07 · Amol Patwardhan

Numerous automatic continuous emotion detection system studies have examined mostly use of videos and images containing individual person expressing emotions. This study examines the detection of spontaneous emotions in …

Edge DetectionEmotion Recognition