paper-with-me

Papers

Kinematic 3D Object Detection in Monocular Video

2020-07-19 · ECCV 2020 8 · Garrick Brazil, Gerard Pons-Moll, Xiaoming Liu, Bernt Schiele

Perceiving the physical world in 3D is fundamental for self-driving applications. Although temporal motion is an invaluable resource to human vision for detection, tracking, and depth perception, such features have not been thoroughly utilized in modern 3D object detectors. In this work, we propose a novel method for monocular video-based 3D object detection which carefully leverages kinematic motion to improve precision of 3D localization. Specifically, we first propose a novel decomposition of object orientation as well as a self-balancing 3D confidence. We show that both components are critical to enable our kinematic model to work effectively. Collectively, using only a single model, we efficiently leverage 3D kinematics from monocular videos to improve the overall localization precision in 3D object detection while also producing useful by-products of scene dynamics (ego-motion and per-object velocity). We achieve state-of-the-art performance on monocular 3D object detection and the Bird's Eye View tasks within the KITTI self-driving dataset.

📄 PDF Abstract BibTeX arXiv:2007.09548

Code (2)

Nicholasli1995/EgoNet pytorch
garrickbrazil/kinematic3d pytorch

Tasks

3D Object DetectionMonocular 3D Object DetectionObjectobject-detectionObject DetectionVehicle Pose Estimation

Similar Papers 제목 키워드 기반

CAMM: Building Category-Agnostic and Animatable 3D Models from Monocular Videos

2023-04-14 · Tianshu Kuai, Akash Karthikeyan, Yash Kant, Ashkan Mirzaei 외

Animating an object in 3D often requires an articulated structure, e.g. a kinematic chain or skeleton of the manipulated object with proper skinning weights, to obtain smooth movements and surface deformations. However, …

ObjectSurface Reconstruction

SurgiPose: Estimating Surgical Tool Kinematics from Monocular Video for Surgical Robot Learning

2025-12-19 · Juo-Tung Chen, XinHao Chen, Ji Woong Kim, Paul Maria Scheikl 외 arxiv

Imitation learning (IL) has shown immense promise in enabling autonomous dexterous manipulation, including learning surgical tasks. To fully unlock the potential of IL for surgery, access to clinical datasets is needed, …

Pose Estimation

Regularizing Dynamic Radiance Fields with Kinematic Fields

2024-07-19 · Woobin Im, Geonho Cha, Sebin Lee, Jumin Lee 외

This paper presents a novel approach for reconstructing dynamic radiance fields from monocular videos. We integrate kinematics with dynamic radiance fields, bridging the gap between the sparse nature of monocular videos …

Kinematics-Driven Gaussian Shape Deformation for Blurry Monocular Dynamic Scenes

2026-05-09 · Yeon-Ji Song, Kiyoung Kwon, Junoh Lee, Jin-Hwa Kim 외 arxiv

Reconstructing dynamic 3D scenes from blurry monocular videos is challenging as motion-induced blur entangles object motion and geometry, hindering geometric consistency. We present Kinematics-GS, a kinematics-aware fram…

Recovering Physically Plausible Human-Object Interactions from Monocular Videos

2026-06-03 · Dingbang Huang, Etienne Vouga, Qixing Huang, Georgios Pavlakos arxiv

In this paper, we propose RePHO, a method to reconstruct physically plausible human-object interactions (HOI) from monocular videos. While existing kinematic-based approaches produce visually plausible motion, they often…

Reinforcement Learning