paper-with-me

홈 › Papers

GeoLanG: Geometry-Aware Language-Guided Grasping with Unified RGB-D Multimodal Learning

2026-02-04 · Rui Tang, Guankun Wang, Long Bai, Huxin Gao, Jiewen Lai, Chi Kit Ng, Jiazheng Wang, Fan Zhang, Hongliang Ren arxiv

Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains highly challenging in cluttered or occluded scenes. Existing methods often rely on multi-stage pipelines that separate object perception and grasping, which leads to limited cross-modal fusion, redundant computation, and poor generalization in cluttered, occluded, or low-texture scenes. To address these limitations, we propose GeoLanG, an end-to-end multi-task framework built upon the CLIP architecture that unifies visual and linguistic inputs into a shared representation space for robust semantic alignment and improved generalization. To enhance target discrimination under occlusion and low-texture conditions, we explore a more effective use of depth information through the Depth-guided Geometric Module (DGGM), which converts depth into explicit geometric priors and injects them into the attention mechanism without additional computational overhead. In addition, we propose Adaptive Dense Channel Integration, which adaptively balances the contributions of multi-layer features to produce more discriminative and generalizable visual representations. Extensive experiments on the OCID-VLG dataset, as well as in both simulation and real-world hardware, demonstrate that GeoLanG enables precise and robust language-guided grasping in complex, cluttered environments, paving the way toward more reliable multimodal robotic manipulation in real-world human-centric settings.

📄 PDF Abstract BibTeX arXiv:2602.04231

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GeoLangBind: Unifying Earth Observation with Agglomerative Vision-Language Foundation Models

2025-03-08 · Zhitong Xiong, Yi Wang, Weikang Yu, Adam J Stewart 외

Earth observation (EO) data, collected from diverse sensors with varying imaging principles, present significant challenges in creating unified analytical frameworks. We present GeoLangBind, a novel agglomerative vision-…

Earth Observation

Learning 6-DOF Grasping Interaction via Deep Geometry-aware 3D Representations

2017-08-24 · Xinchen Yan, Jasmine Hsu, Mohi Khansari, Yunfei Bai 외

This paper focuses on the problem of learning 6-DOF grasping with a parallel jaw gripper in simulation. We propose the notion of a geometry-aware representation in grasping based on the assumption that knowledge of 3D ge…

3D geometry3D Geometry Prediction3D Shape ModelingData Augmentation

Learning 6-DoF Fine-grained Grasp Detection Based on Part Affordance Grounding

2023-01-27 · Yaoxian Song, Penglei Sun, Piaopiao Jin, Yi Ren 외

Robotic grasping is a fundamental ability for a robot to interact with the environment. Current methods focus on how to obtain a stable and reliable grasping pose in object level, while little work has been studied on pa…

3D geometryRepresentation LearningRobotic Grasping

Language-Guided Grasping under Partial Observation for Mobile Manipulation in Field Inspection and Maintenance

2026-03-09 · Dilermando Almeida, Juliano Negri, Guilherme Lazzarini, Thiago H. Segreto 외 arxiv

Offshore inspection and maintenance have increasingly been using legged robots for routine sensing, yet many useful interventions still require physical interaction with tools, containers, and task-relevant objects. Empl…

GAPG: Geometry Aware Push-Grasping Synergy for Goal-Oriented Manipulation in Clutter

2026-03-22 · Lijingze Xiao, Jinhong Du, Yang Cong, Supeng Diao 외 arxiv

Grasping target objects is a fundamental skill for robotic manipulation, but in cluttered environments with stacked or occluded objects, a single-step grasp is often insufficient. To address this, previous work has intro…