Objectron: A Large Scale Dataset of Object-Centric Videos in the Wild with Pose Annotations
3D object detection has recently become popular due to many applications in robotics, augmented reality, autonomy, and image retrieval. We introduce the Objectron dataset to advance the state of the art in 3D object detection and foster new research and applications, such as 3D object tracking, view synthesis, and improved 3D shape representation. The dataset contains object-centric short videos with pose annotations for nine categories and includes 4 million annotated images in 14,819 annotated videos. We also propose a new evaluation metric, 3D Intersection over Union, for 3D object detection. We demonstrate the usefulness of our dataset in 3D object detection tasks by providing baseline models trained on this dataset. Our dataset and evaluation source code are available online at http://www.objectron.dev
Code (1)
Tasks
3D Object Detection3D Object Tracking3D Shape RepresentationImage RetrievalMonocular 3D Object DetectionObjectobject-detectionObject DetectionObject TrackingRetrievalMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Keypoint-Based Category-Level Object Pose Tracking from an RGB Sequence with Uncertainty Estimation
We propose a single-stage, category-level 6-DoF pose estimation algorithm that simultaneously detects and tracks instances of objects within a known category. Our method takes as input the previous and current frame from…
Pose EstimationPose TrackingEgoObjects: A Large-Scale Egocentric Dataset for Fine-Grained Object Understanding
Object understanding in egocentric visual data is arguably a fundamental research topic in egocentric vision. However, existing object datasets are either non-egocentric or have limitations in object categories, visual c…
Continual LearningObjectobject-detectionObject DetectionDeveloping Vision-Language-Action Model from Egocentric Videos
Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-La…
Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
Learning to use tools or objects in common scenes, particularly handling them in various ways as instructed, is a key challenge for developing interactive robots. Training models to generate such manipulation traject…
validThinking in Boxes: 3D Editing in Real Images Made Easy
Text and 2D-conditioning interfaces provide weak, ambiguous control over spatial transformations in image editing -- particularly under large object motions and camera changes. Prior work has used 3D primitives such as b…
Image Editing