paper-with-me

홈 › Papers

Geometry-Conditioned Diffusion for Occlusion-Robust In-Bed Pose Estimation

2026-04-26 · Navid Aslankhani Khameneh, Marco Carletti, Cigdem Beyan arxiv

Robust in-bed human pose estimation under blanket occlusion remains challenging due to the scarcity of reliable labeled training data for heavily covered poses. Existing approaches rely on multi-modal sensing or image-to-image translation frameworks that remain conditioned on visible source imagery, limiting scalability and pose diversity. In this work, we reformulate occlusion-aware augmentation as a geometry-conditioned generative modeling task. We conduct a systematic comparison of deterministic masking, unpaired translation, paired diffusion-based translation, and a proposed pose-conditioned Latent Diffusion Model (Pose-LDM). Unlike image-guided methods, Pose-LDM synthesizes blanket-covered images directly from skeletal keypoints, eliminating dependence on paired supervision and pixel-level source-image conditioning while enabling generation from arbitrary pose inputs. All augmentation strategies are evaluated through their impact on downstream pose estimation under a fixed backbone. Pose- LDM achieves the highest strict localization accuracy under severe occlusion while maintaining overall detection performance comparable to paired diffusion models, approaching the performance of fully supervised training. These results demonstrate that geometry-conditioned diffusion provides an effective and supervision-efficient pathway toward occlusion-robust inbed pose estimation without modifying the sensing pipeline. The code is available at: github.com/navidTerraNova/ GeoDiffPose.

📄 PDF Abstract BibTeX arXiv:2604.23651

Code (0)

등록된 구현이 없습니다.

Tasks

Image-to-Image TranslationPose Estimation

Similar Papers 제목 키워드 기반

StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space

2025-12-11 · Tjark Behrens, Anton Obukhov, Bingxin Ke, Fabio Tosi 외 arxiv

We introduce StereoSpace, a diffusion-based framework for monocular-to-stereo synthesis that models geometry purely through viewpoint conditioning, without explicit depth or warping. A canonical rectified space and the c…

DiffPose: Toward More Reliable 3D Pose Estimation

2022-11-30 · CVPR 2023 1 · Jia Gong, Lin Geng Foo, Zhipeng Fan, Qiuhong Ke 외

Monocular 3D human pose estimation is quite challenging due to the inherent ambiguity and occlusion, which often lead to high uncertainty and indeterminacy. On the other hand, diffusion models have recently emerged as an…

3D Human Pose Estimation3D Pose EstimationMonocular 3D Human Pose EstimationPose Estimation

Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild

2026-06-22 · Yehonathan Litman, Xiaoxuan Ma, Manan Shah, Nicolas Ugrinovic 외 arxiv

Reconstructing dynamic non-rigid objects from monocular video requires integrating visual cues from direct observations with data-driven priors over geometry and appearance. Prior approaches either learn to directly pred…

Single-View 3D Reconstruction

Video Generative Models as Geometry Learner

2026-08-28 · Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu 외 arxiv

Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as image-conditioned generation. Leveraging off-the-shelf image diffusion models, they either (i) train task-…

SG-DOR: Learning Scene Graphs with Direction-Conditioned Occlusion Reasoning for Pepper Plants

2026-03-06 · Rohit Menon, Niklas Mueller-Goldingen, Sicong Pan, Gokul Krishna Chenchani 외 arxiv

Robotic harvesting in dense crop canopies requires effective interventions that depend not only on geometry, but also on explicit, direction-conditioned relations identifying which organs obstruct a target fruit. We pres…

Point Clouds