paper-with-me

Papers

Interactive Visual Task Learning for Robots

2023-12-20 · Weiwei Gu, Anant Sah, Nakul Gopalan

We present a framework for robots to learn novel visual concepts and tasks via in-situ linguistic interactions with human users. Previous approaches have either used large pre-trained visual models to infer novel objects zero-shot, or added novel concepts along with their attributes and representations to a concept hierarchy. We extend the approaches that focus on learning visual concept hierarchies by enabling them to learn novel concepts and solve unseen robotics tasks with them. To enable a visual concept learner to solve robotics tasks one-shot, we developed two distinct techniques. Firstly, we propose a novel approach, Hi-Viscont(HIerarchical VISual CONcept learner for Task), which augments information of a novel concept to its parent nodes within a concept hierarchy. This information propagation allows all concepts in a hierarchy to update as novel concepts are taught in a continual learning setting. Secondly, we represent a visual task as a scene graph with language annotations, allowing us to create novel permutations of a demonstrated task zero-shot in-situ. We present two sets of results. Firstly, we compare Hi-Viscont with the baseline model (FALCON) on visual question answering(VQA) in three domains. While being comparable to the baseline model on leaf level concepts, Hi-Viscont achieves an improvement of over 9% on non-leaf concepts on average. We compare our model's performance against the baseline FALCON model. Our framework achieves 33% improvements in success rate metric, and 19% improvements in the object level accuracy compared to the baseline model. With both of these results we demonstrate the ability of our model to learn tasks and concepts in a continual learning setting on the robot.

📄 PDF Abstract BibTeX arXiv:2312.13219

Code (0)

등록된 구현이 없습니다.

Tasks

Continual LearningNovel ConceptsQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Focus 설명 없음
+ ( 1 ) ⟷ 888 ⟷ ( 829 ) ⟷ 0881||How do I resolve a dispute on Expedia? How do I resolve a dispute on Expedia contact their support at + ( 1 ) ⟷ 888 ⟷ ( 829 ) ⟷ 0881 or + ( 1 ) ⟷ 805 ⟷ ( 330 ) ⟷ 4056. Provide booking details and explain the issue…

Similar Papers 제목 키워드 기반

Building an Affordances Map with Interactive Perception

2019-03-11 · Leni K. Le Goff, Oussama Yaakoubi, Alexandre Coninx, Stephane Doncieux

Robots need to understand their environment to perform their task. If it is possible to pre-program a visual scene analysis process in closed environments, robots operating in an open environment would benefit from the a…

General ClassificationScene Understanding

Enabling Robots to Draw and Tell: Towards Visually Grounded Multimodal Description Generation

2021-01-14 · Ting Han, Sina Zarrieß

Socially competent robots should be equipped with the ability to perceive the world that surrounds them and communicate about it in a human-like manner. Representative skills that exhibit such ability include generating …

Interactive Robot Programming for Surface Finishing via Task-Centric Mixed Reality Interfaces

2025-12-29 · Christoph Willibald, Lugh Martensen, Thomas Eiband, Dongheui Lee arxiv

Lengthy setup processes that require robotics expertise remain a major barrier to deploying robots for tasks involving high product variability and small batch sizes. As a result, collaborative robots, despite their adva…

RoboLinker: A Diffusion-model-based Matching Clothing Generator Between Humans and Companion Robots

2025-08-02 · Jing Tang, Qing Xiao, Kunxu Du, Zaiqiao Ye arxiv

We present RoboLinker, a generative design system that creates matching outfits for humans and their robots. Using a diffusion-based model, the system takes a robot image and a style prompt from users as input, and outpu…

EgoPriMo: Egocentric Motion Generation for Interactive Humanoid Control

2026-06-07 · Haoyang Ge, Peng Ren, Yukun Shi, Cong Huang 외 arxiv

Humanoid robots require whole-body motions that adapt to scene context, task requirements, and user intent. Motion tracking reproduces specified trajectories, and humanoid vision-language-action systems provide semantic …