paper-with-me

Papers

Exploiting Spatial Invariance for Scalable Unsupervised Object Tracking

2019-11-20 · Eric Crawford, Joelle Pineau

The ability to detect and track objects in the visual world is a crucial skill for any intelligent agent, as it is a necessary precursor to any object-level reasoning process. Moreover, it is important that agents learn to track objects without supervision (i.e. without access to annotated training videos) since this will allow agents to begin operating in new environments with minimal human assistance. The task of learning to discover and track objects in videos, which we call \textit{unsupervised object tracking}, has grown in prominence in recent years; however, most architectures that address it still struggle to deal with large scenes containing many objects. In the current work, we propose an architecture that scales well to the large-scene, many-object setting by employing spatially invariant computations (convolutions and spatial attention) and representations (a spatially local object specification scheme). In a series of experiments, we demonstrate a number of attractive features of our architecture; most notably, that it outperforms competing methods at tracking objects in cluttered scenes with many objects, and that it can generalize well to videos that are larger and/or contain more objects than videos encountered during training.

📄 PDF Abstract BibTeX arXiv:1911.09033

Code (1)

e2crawfo/silot 공식 구현 tf

Tasks

ObjectObject Tracking

Similar Papers 제목 키워드 기반

Hierarchically Decoupled Spatial-Temporal Contrast for Self-supervised Video Representation Learning

2020-11-23 · Zehua Zhang, David Crandall

We present a novel technique for self-supervised video representation learning by: (a) decoupling the learning objective into two contrastive subtasks respectively emphasizing spatial and temporal features, and (b) perfo…

Action RecognitionContrastive LearningRepresentation Learning

Unsupervised Part-Based Disentangling of Object Shape and Appearance

2019-03-16 · CVPR 2019 6 · Dominik Lorenz, Leonard Bereska, Timo Milbich, Björn Ommer

Large intra-class variation is the result of changes in multiple object characteristics. Images, however, only show the superposition of different variable factors such as appearance or shape. Therefore, learning to dise…

Appearance TransferImage GenerationObjectPose Prediction+3

Deformable Part-based Fully Convolutional Network for Object Detection

2017-07-19 · Taylor Mordan, Nicolas Thome, Matthieu Cord, Gilles Henaff

Existing region-based object detectors are limited to regions with fixed box geometry to represent objects, even if those are highly non-rectangular. In this paper we introduce DP-FCN, a deep model for object detection w…

Objectobject-detectionObject Detection

Spatially Invariant Unsupervised 3D Object-Centric Learning and Scene Decomposition

2021-06-10 · Tianyu Wang, Miaomiao Liu, Kee Siong Ng

We tackle the problem of object-centric learning on point clouds, which is crucial for high-level relational reasoning and scalable machine intelligence. In particular, we introduce a framework, SPAIR3D, to factorize a 3…

ObjectRelational ReasoningSemantic Segmentation

Rethinking Spatial Invariance of Convolutional Networks for Object Counting

2022-06-10 · CVPR 2022 1 · Zhi-Qi Cheng, Qi Dai, Hong Li, Jingkuan Song 외

Previous work generally believes that improving the spatial invariance of convolutional networks is the key to object counting. However, after verifying several mainstream counting networks, we surprisingly found too str…

Crowd CountingObjectObject CountingPosition