paper-with-me

Papers

Learning Articulated Motion Models from Visual and Lingual Signals

2015-11-17 · Zhengyang Wu, Mohit Bansal, Matthew R. Walter

In order for robots to operate effectively in homes and workplaces, they must be able to manipulate the articulated objects common within environments built for and by humans. Previous work learns kinematic models that prescribe this manipulation from visual demonstrations. Lingual signals, such as natural language descriptions and instructions, offer a complementary means of conveying knowledge of such manipulation models and are suitable to a wide range of interactions (e.g., remote manipulation). In this paper, we present a multimodal learning framework that incorporates both visual and lingual information to estimate the structure and parameters that define kinematic models of articulated objects. The visual signal takes the form of an RGB-D image stream that opportunistically captures object motion in an unprepared scene. Accompanying natural language descriptions of the motion constitute the lingual signal. We present a probabilistic language model that uses word embeddings to associate lingual verbs with their corresponding kinematic structures. By exploiting the complementary nature of the visual and lingual input, our method infers correct kinematic structures for various multiple-part objects on which the previous state-of-the-art, visual-only system fails. We evaluate our multimodal learning framework on a dataset comprised of a variety of household objects, and demonstrate a 36% improvement in model accuracy over the vision-only baseline.

📄 PDF Abstract BibTeX arXiv:1511.05526

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingWord Embeddings

Similar Papers 제목 키워드 기반

AirDOS: Dynamic SLAM benefits from Articulated Objects

2021-09-21 · Yuheng Qiu, Chen Wang, Wenshan Wang, Mina Henein 외

Dynamic Object-aware SLAM (DOS) exploits object-level information to enable robust motion estimation in dynamic environments. Existing methods mainly focus on identifying and excluding dynamic objects from the optimizati…

Camera Pose EstimationMotion EstimationObjectPose Estimation

ArticulatedGS: Self-supervised Digital Twin Modeling of Articulated Objects using 3D Gaussian Splatting

2025-01-01 · CVPR 2025 1 · Junfu Guo, Yu Xin, Gaoyi Liu, Kai Xu 외

We tackle the challenge of concurrent reconstruction at the part level with the RGB appearance and estimation of motion parameters for building digital twins of articulated objects using the 3D Gaussian Splatting (3D…

Motion Estimation

On the role of Lip Articulation in Visual Speech Perception

2022-03-18 · Zakaria Aldeneh, Masha Fedzechkina, Skyler Seto, Katherine Metcalf 외

Generating realistic lip motion from audio to simulate speech production is critical for driving natural character animation. Previous research has shown that traditional metrics used to optimize and assess models for ge…

MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction

2026-03-19 · Haitian Li, Haozhe Xie, Junxiang Xu, Beichen Wen 외 arxiv

Reconstructing articulated 3D objects from a single image requires jointly inferring object geometry, part structure, and motion parameters from limited visual evidence. A key difficulty lies in the entanglement between …

3D ReconstructionVideo Generation

From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation

2026-05-09 · Bohan Li, Shuojue Yang, Baorui Peng, Xianda Guo 외 arxiv

Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimensional control vectors must precisely govern complex image-space evoluti…

Domain GeneralizationVideo GenerationPose Tracking