paper-with-me

Papers

D-SCo: Dual-Stream Conditional Diffusion for Monocular Hand-Held Object Reconstruction

2023-11-23 · Bowen Fu, Gu Wang, Chenyangguang Zhang, Yan Di, Ziqin Huang, Zhiying Leng, Fabian Manhardt, Xiangyang Ji, Federico Tombari

Reconstructing hand-held objects from a single RGB image is a challenging task in computer vision. In contrast to prior works that utilize deterministic modeling paradigms, we employ a point cloud denoising diffusion model to account for the probabilistic nature of this problem. In the core, we introduce centroid-fixed dual-stream conditional diffusion for monocular hand-held object reconstruction (D-SCo), tackling two predominant challenges. First, to avoid the object centroid from deviating, we utilize a novel hand-constrained centroid fixing paradigm, enhancing the stability of diffusion and reverse processes and the precision of feature projection. Second, we introduce a dual-stream denoiser to semantically and geometrically model hand-object interactions with a novel unified hand-object semantic embedding, enhancing the reconstruction performance of the hand-occluded region of the object. Experiments on the synthetic ObMan dataset and three real-world datasets HO3D, MOW and DexYCB demonstrate that our approach can surpass all other state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2311.14189

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingObjectObject Reconstruction

Methods 이 논문이 사용한 방법론

Focus 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models

2025-03-07 · Mark YU, WenBo Hu, Jinbo Xing, Ying Shan

We present TrajectoryCrafter, a novel approach to redirect camera trajectories for monocular videos. By disentangling deterministic view transformations from stochastic content generation, our method achieves precise con…

InterHandGen: Two-Hand Interaction Generation via Cascaded Reverse Diffusion

2024-03-26 · CVPR 2024 1 · Jihyun Lee, Shunsuke Saito, Giljoo Nam, Minhyuk Sung 외

We present InterHandGen, a novel framework that learns the generative prior of two-hand interaction. Sampling from our model yields plausible and diverse two-hand shapes in close interaction with or without an object. Ou…

Diversity

DiffHand: End-to-End Hand Mesh Reconstruction via Diffusion Models

2023-05-23 · Lijun Li, Li'an Zhuo, Bang Zhang, Liefeng Bo 외

Hand mesh reconstruction from the monocular image is a challenging task due to its depth ambiguity and severe occlusion, there remains a non-unique mapping between the monocular image and hand mesh. To address this, we d…

DecoderDenoising

Marigold-DC: Zero-Shot Monocular Depth Completion with Guided Diffusion

2024-12-18 · Massimiliano Viola, Kevin Qu, Nando Metzger, Bingxin Ke 외

Depth completion upgrades sparse depth measurements into dense depth maps guided by a conventional image. Existing methods for this highly ill-posed task operate in tightly constrained settings and tend to struggle when …

DenoisingDepth CompletionDepth EstimationMonocular Depth Estimation+1

Veila: Panoramic LiDAR Generation from a Monocular RGB Image

2025-08-05 · Youquan Liu, Lingdong Kong, Weidong Yang, Ao Liang 외 arxiv

Realistic and controllable panoramic LiDAR data generation is critical for scalable 3D perception in autonomous driving and robotics. Existing methods either perform unconditional generation with poor controllability or …

LIDAR Semantic SegmentationAutonomous DrivingData Augmentation