paper-with-me

Papers

Objectron: A Large Scale Dataset of Object-Centric Videos in the Wild with Pose Annotations

2020-12-18 · CVPR 2021 1 · Adel Ahmadyan, Liangkai Zhang, Jianing Wei, Artsiom Ablavatski, Matthias Grundmann

3D object detection has recently become popular due to many applications in robotics, augmented reality, autonomy, and image retrieval. We introduce the Objectron dataset to advance the state of the art in 3D object detection and foster new research and applications, such as 3D object tracking, view synthesis, and improved 3D shape representation. The dataset contains object-centric short videos with pose annotations for nine categories and includes 4 million annotated images in 14,819 annotated videos. We also propose a new evaluation metric, 3D Intersection over Union, for 3D object detection. We demonstrate the usefulness of our dataset in 3D object detection tasks by providing baseline models trained on this dataset. Our dataset and evaluation source code are available online at http://www.objectron.dev

📄 PDF Abstract BibTeX arXiv:2012.09988

Code (1)

google-research-datasets/Objectron 공식 구현 tf

Tasks

3D Object Detection3D Object Tracking3D Shape RepresentationImage RetrievalMonocular 3D Object DetectionObjectobject-detectionObject DetectionObject TrackingRetrieval

Methods 이 논문이 사용한 방법론

Pointwise Convolution Pointwise Convolution is a type of convolution that uses a 1x1 kernel: a kernel that iterates through every single point. This…
Depthwise Convolution Depthwise Convolution is a type of convolution where we apply a single convolutional filter for each input channel. In the regular 2D…
Sigmoid Activation 설명 없음
Depthwise Separable Convolution While standard convolution performs the channelwise and spatial-wise computation in one step, Depthwise Separable Convolution
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Batch Normalization 설명 없음
RMSProp RMSProp is an unpublished adaptive learning rate optimizer proposed by Geoff Hinton. The motivation…
Squeeze-and-Excitation Block The Squeeze-and-Excitation Block is an architectural unit designed to improve the representational power of a network by enabling it to perform dynamic channel-wise feature…

Similar Papers 제목 키워드 기반

Keypoint-Based Category-Level Object Pose Tracking from an RGB Sequence with Uncertainty Estimation

2022-05-23 · Yunzhi Lin, Jonathan Tremblay, Stephen Tyree, Patricio A. Vela 외

We propose a single-stage, category-level 6-DoF pose estimation algorithm that simultaneously detects and tracks instances of objects within a known category. Our method takes as input the previous and current frame from…

Pose EstimationPose Tracking

EgoObjects: A Large-Scale Egocentric Dataset for Fine-Grained Object Understanding

2023-09-15 · ICCV 2023 1 · Chenchen Zhu, Fanyi Xiao, Andres Alvarado, Yasmine Babaei 외

Object understanding in egocentric visual data is arguably a fundamental research topic in egocentric vision. However, existing object datasets are either non-egocentric or have limitations in object categories, visual c…

Continual LearningObjectobject-detectionObject Detection

Developing Vision-Language-Action Model from Egocentric Videos

2025-09-26 · Tomoya Yoshida, Shuhei Kurita, Taichi Nishimura, Shinsuke Mori arxiv

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-La…

Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision

2025-01-01 · CVPR 2025 1 · Tomoya Yoshida, Shuhei Kurita, Taichi Nishimura, Shinsuke Mori

Learning to use tools or objects in common scenes, particularly handling them in various ways as instructed, is a key challenge for developing interactive robots. Training models to generate such manipulation traject…

valid

Thinking in Boxes: 3D Editing in Real Images Made Easy

2026-06-18 · Pradhaan S Bhat, Naveen Chandra R, Rishubh Parihar, Vaibhav Vavilala 외 arxiv

Text and 2D-conditioning interfaces provide weak, ambiguous control over spatial transformations in image editing -- particularly under large object motions and camera changes. Prior work has used 3D primitives such as b…

Image Editing