paper-with-me

홈 › Papers

Semantically Grounded Object Matching for Robust Robotic Scene Rearrangement

2021-11-15 · Walter Goodwin, Sagar Vaze, Ioannis Havoutis, Ingmar Posner

Object rearrangement has recently emerged as a key competency in robot manipulation, with practical solutions generally involving object detection, recognition, grasping and high-level planning. Goal-images describing a desired scene configuration are a promising and increasingly used mode of instruction. A key outstanding challenge is the accurate inference of matches between objects in front of a robot, and those seen in a provided goal image, where recent works have struggled in the absence of object-specific training data. In this work, we explore the deterioration of existing methods' ability to infer matches between objects as the visual shift between observed and goal scenes increases. We find that a fundamental limitation of the current setting is that source and target images must contain the same $\textit{instance}$ of every object, which restricts practical deployment. We present a novel approach to object matching that uses a large pre-trained vision-language model to match objects in a cross-instance setting by leveraging semantics together with visual features as a more robust, and much more general, measure of similarity. We demonstrate that this provides considerably improved matching performance in cross-instance settings, and can be used to guide multi-object rearrangement with a robot manipulator from an image that shares no object $\textit{instances}$ with the robot's scene.

📄 PDF Abstract BibTeX arXiv:2111.07975

Code (1)

applied-ai-lab/object_matching 공식 구현

Tasks

Language ModellingObjectobject-detectionObject DetectionObject RearrangementRobot Manipulation

Similar Papers 제목 키워드 기반

Analogical Trajectory Transfer

2026-05-14 · Junho Kim, Eun Sun Lee, Gwangtak Bae, Seunggu Kang 외 arxiv

We study analogical trajectory transfer, where the goal is to translate motion trajectories in one 3D environment to a semantically analogous location in another. Such a capacity would enable machines to perform analogic…

Spatial ReasoningGraph Matching

FATE: Closed-Loop Feasibility-Aware Task Generation with Active Repair for Physically Grounded Robotic Curricula

2026-03-02 · Bingchuan Wei, Bingqi Huang, Jingheng Ma, Zeyu zhang 외 arxiv

Recent breakthroughs in generative simulation have harnessed Large Language Models (LLMs) to generate diverse robotic task curricula, yet these open-loop paradigms frequently produce linguistically coherent but physicall…

FlowHOI: Flow-based Semantics-Grounded Generation of Hand-Object Interactions for Dexterous Robot Manipulation

2026-02-13 · Huajian Zeng, Lingyun Chen, Jiaqi Yang, Yuantai Zhang 외 arxiv

Recent vision-language-action (VLA) models can generate plausible end-effector motions, yet they often fail in long-horizon, contact-rich tasks because the underlying hand-object interaction (HOI) structure is not explic…

Robot ManipulationAction Recognition

CityRAG: Stepping Into a City via Spatially-Grounded Video Generation

2026-04-21 · Gene Chou, Charles Herrmann, Kyle Genova, Boyang Deng 외 arxiv

We address the problem of generating a 3D-consistent, navigable environment that is spatially grounded: a simulation of a real location. Existing video generative models can produce a plausible sequence that is consisten…

Autonomous DrivingVideo Generation

Scene-agnostic Hierarchical Bimanual Task Planning via Visual Affordance Reasoning

2025-12-10 · Kwang Bin Lee, Jiho Kang, Sung-Hee Lee arxiv

Embodied agents operating in open environments must translate high-level instructions into grounded, executable behaviors, often requiring coordinated use of both hands. While recent foundation models offer strong semant…