paper-with-me

홈 › Papers

CLIPUNetr: Assisting Human-robot Interface for Uncalibrated Visual Servoing Control with CLIP-driven Referring Expression Segmentation

2023-09-17 · Chen Jiang, Yuchen Yang, Martin Jagersand

The classical human-robot interface in uncalibrated image-based visual servoing (UIBVS) relies on either human annotations or semantic segmentation with categorical labels. Both methods fail to match natural human communication and convey rich semantics in manipulation tasks as effectively as natural language expressions. In this paper, we tackle this problem by using referring expression segmentation, which is a prompt-based approach, to provide more in-depth information for robot perception. To generate high-quality segmentation predictions from referring expressions, we propose CLIPUNetr - a new CLIP-driven referring expression segmentation network. CLIPUNetr leverages CLIP's strong vision-language representations to segment regions from referring expressions, while utilizing its ``U-shaped'' encoder-decoder architecture to generate predictions with sharper boundaries and finer structures. Furthermore, we propose a new pipeline to integrate CLIPUNetr into UIBVS and apply it to control robots in real-world environments. In experiments, our method improves boundary and structure measurements by an average of 120% and can successfully assist real-world UIBVS control in an unstructured manipulation environment.

📄 PDF Abstract BibTeX arXiv:2309.09183

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderReferring ExpressionReferring Expression SegmentationSegmentationSemantic Segmentation

Methods 이 논문이 사용한 방법론

fail 설명 없음

Similar Papers 제목 키워드 기반

D-RMGPT: Robot-assisted collaborative tasks driven by large multimodal models

2024-08-21 · M. Forlini, M. Babcinschi, G. Palmieri, P. Neto

Collaborative robots are increasingly popular for assisting humans at work and daily tasks. However, designing and setting up interfaces for human-robot collaboration is challenging, requiring the integration of multiple…

Age-Related Differences in the Perception of Eye-Gaze from a Social Robot

2026-03-09 · Lucas Morillo-Mendez, Martien G. S. Schrooten, Oscar Martinez Mozos arxiv

There is an increasing interest in social robots assisting older adults during daily life tasks. In this context, non-verbal cues such as deictic gaze are important in natural communication in human-robot interaction. Ho…

Explainable OOHRI: Communicating Robot Capabilities and Limitations as Augmented Reality Affordances

2026-01-21 · Lauren W. Wang, Mohamed Kari, Parastoo Abtahi arxiv

Human interaction is essential for issuing personalized instructions and assisting robots when failure is likely. However, robots remain largely black boxes, offering users little insight into their evolving capabilities…

Explanation Generation

Analyzing Reluctance to Ask for Help When Cooperating With Robots: Insights to Integrate Artificial Agents in HRC

2025-09-01 · Ane San Martin, Michael Hagenow, Julie Shah, Johan Kildal 외 arxiv

As robot technology advances, collaboration between humans and robots will become more prevalent in industrial tasks. When humans run into issues in such scenarios, a likely future involves relying on artificial agents o…

Graph2Bots, Unsupervised Assistance for Designing Chatbots

2019-09-01 · WS 2019 9 · Jean-Leon Bouraoui, Sonia Le Meitour, Romain Carbou, Lina M. Rojas Barahona 외

We present Graph2Bots, a tool for assisting conversational agent designers. It extracts a graph representation from human-human conversations by using unsupervised learning. The generated graph contains the main stages o…