paper-with-me

Papers

Object-Scene-Camera Decomposition and Recomposition for Data-Efficient Monocular 3D Object Detection

2026-02-24 · Zhaonian Kuang, Rui Ding, Meng Yang, Xinhu Zheng, Gang Hua arxiv

Monocular 3D object detection (M3OD) is intrinsically ill-posed, hence training a high-performance deep learning based M3OD model requires a humongous amount of labeled data with complicated visual variation from diverse scenes, variety of objects and camera poses.However, we observe that, due to strong human bias, the three independent entities, i.e., object, scene, and camera pose, are always tightly entangled when an image is captured to construct training data. More specifically, specific 3D objects are always captured in particular scenes with fixed camera poses, and hence lacks necessary diversity. Such tight entanglement induces the challenging issues of insufficient utilization and overfitting to uniform training data. To mitigate this, we propose an online object-scene-camera decomposition and recomposition data manipulation scheme to more efficiently exploit the training data. We first fully decompose training images into textured 3D object point models and background scenes in an efficient computation and storage manner. We then continuously recompose new training images in each epoch by inserting the 3D objects into the freespace of the background scenes, and rendering them with perturbed camera poses from textured 3D point representation. In this way, the refreshed training data in all epochs can cover the full spectrum of independent object, scene, and camera pose combinations. This scheme can serve as a plug-and-play component to boost M3OD models, working flexibly with both fully and sparsely supervised settings. In the sparsely-supervised setting, objects closest to the ego-camera for all instances are sparsely annotated. We then can flexibly increase the annotated objects to control annotation cost. For validation, our method is widely applied to five representative M3OD models and evaluated on both the KITTI and the more complicated Waymo datasets.

📄 PDF Abstract BibTeX arXiv:2602.20627

Code (0)

등록된 구현이 없습니다.

Tasks

Monocular 3D Object Detection

Similar Papers 제목 키워드 기반

Frequency-Domain Decomposition and Recomposition for Robust Audio-Visual Segmentation

2025-09-23 · Yunzhe Shen, Kai Peng, Leiye Liu, Wei Ji 외 arxiv

Audio-visual segmentation (AVS) plays a critical role in multimodal machine learning by effectively integrating audio and visual cues to precisely segment objects or regions within visual scenes. Recent AVS methods have …

Vision-Aided Frame-Capture-Based CSI Recomposition for WiFi Sensing: A Multimodal Approach

2022-06-03 · Hiroki Shimomura, Yusuke Koda, Takamochi Kanda, Koji Yamamoto 외

Recompositing channel state information (CSI) from the beamforming feedback matrix (BFM), which is a compressed version of CSI and can be captured because of its lack of encryption, is an alternative way of implementing …

Multimodal Deep Learning

Vista4D: Video Reshooting with 4D Point Clouds

2026-04-23 · Kuan Heng Lin, Zhizheng Liu, Pablo Salamanca, Yash Kant 외 arxiv

We present Vista4D, a robust and flexible video reshooting framework that grounds the input video and target cameras in a 4D point cloud. Specifically, given an input video, our method re-synthesizes the scene with the s…

Depth EstimationPoint Clouds

CompoNeRF: Text-guided Multi-object Compositional NeRF with Editable 3D Scene Layout

2023-03-24 · Haotian Bai, Yuanhuiyi Lyu, Lutao Jiang, Sijia Li 외

Text-to-3D form plays a crucial role in creating editable 3D scenes for AR/VR. Recent advances have shown promise in merging neural radiance fields (NeRFs) with pre-trained diffusion models for text-to-3D object generati…

NeRFObjectScene GenerationText to 3D

Proactive Scene Decomposition and Reconstruction

2025-10-17 · Baicheng Li, Zike Yan, Dong Wu, Hongbin Zha arxiv

Human behaviors are the major causes of scene dynamics and inherently contain rich cues regarding the dynamics. This paper formalizes a new task of proactive scene decomposition and reconstruction, an online approach tha…

Pose Estimation