paper-with-me

홈 › Papers

From Words to Poses: Enhancing Novel Object Pose Estimation with Vision Language Models

2024-09-09 · Tessa Pulli, Stefan Thalhammer, Simon Schwaiger, Markus Vincze

Robots are increasingly envisioned to interact in real-world scenarios, where they must continuously adapt to new situations. To detect and grasp novel objects, zero-shot pose estimators determine poses without prior knowledge. Recently, vision language models (VLMs) have shown considerable advances in robotics applications by establishing an understanding between language input and image input. In our work, we take advantage of VLMs zero-shot capabilities and translate this ability to 6D object pose estimation. We propose a novel framework for promptable zero-shot 6D object pose estimation using language embeddings. The idea is to derive a coarse location of an object based on the relevancy map of a language-embedded NeRF reconstruction and to compute the pose estimate with a point cloud registration method. Additionally, we provide an analysis of LERF's suitability for open-set object pose estimation. We examine hyperparameters, such as activation thresholds for relevancy maps and investigate the zero-shot capabilities on an instance- and category-level. Furthermore, we plan to conduct robotic grasping experiments in a real-world setting.

📄 PDF Abstract BibTeX arXiv:2409.05413

Code (0)

등록된 구현이 없습니다.

Tasks

6D Pose Estimation using RGBNeRFObjectPoint Cloud RegistrationPose EstimationRobotic Grasping

Similar Papers 제목 키워드 기반

UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References

2025-06-09 · CVPR 2025 1 · Ming-Feng Li, Xin Yang, Fu-En Wang, Hritam Basak 외

6D object pose estimation has shown strong generalizability to novel objects. However, existing methods often require either a complete, well-reconstructed 3D model or numerous reference images that fully cover the objec…

6D Pose Estimation using RGBImage to 3DObjectPose Estimation

BoxDreamer: Dreaming Box Corners for Generalizable Object Pose Estimation

2025-04-10 · Yuanhong Yu, Xingyi He, Chen Zhao, Junhao Yu 외

This paper presents a generalizable RGB-based approach for object pose estimation, specifically designed to address challenges in sparse-view settings. While existing methods can estimate the poses of unseen objects, the…

ObjectPose Estimation

Efficient Online 3D Multi-Camera Multi-Object Tracking and Pose Estimation

2026-04-16 · Linh Van Ma, Tran Thien Dat Nguyen, Juhua Hu, Wei Cheng 외 arxiv

This paper proposes a fast and online method for jointly performing 3D multi-object tracking and pose estimation using multiple monocular cameras. Our algorithm requires only 2D bounding box and pose detections, eliminat…

3D Multi-Object TrackingComputational EfficiencyPose Estimation

Counterfactual Samples Constructing and Training for Commonsense Statements Estimation

2024-12-29 · Chong Liu, Zaiwen Feng, Lin Liu, Zhenyun Deng 외

Plausibility Estimation (PE) plays a crucial role for enabling language models to objectively comprehend the real world. While large language models (LLMs) demonstrate remarkable capabilities in PE tasks but sometimes pr…

counterfactualSentence

Retraining-free Customized ASR for Enharmonic Words Based on a Named-Entity-Aware Model and Phoneme Similarity Estimation

2023-05-29 · Yui Sudo, Kazuya Hata, Kazuhiro Nakadai

End-to-end automatic speech recognition (E2E-ASR) has the potential to improve performance, but a specific issue that needs to be addressed is the difficulty it has in handling enharmonic words: named entities (NEs) with…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition