paper-with-me

홈 › Papers

Semantic-Based Active Perception for Humanoid Visual Tasks with Foveal Sensors

2024-04-16 · João Luzio, Alexandre Bernardino, Plinio Moreno

The aim of this work is to establish how accurately a recent semantic-based foveal active perception model is able to complete visual tasks that are regularly performed by humans, namely, scene exploration and visual search. This model exploits the ability of current object detectors to localize and classify a large number of object classes and to update a semantic description of a scene across multiple fixations. It has been used previously in scene exploration tasks. In this paper, we revisit the model and extend its application to visual search tasks. To illustrate the benefits of using semantic information in scene exploration and visual search tasks, we compare its performance against traditional saliency-based models. In the task of scene exploration, the semantic-based method demonstrates superior performance compared to the traditional saliency-based model in accurately representing the semantic information present in the visual scene. In visual search experiments, searching for instances of a target class in a visual field containing multiple distractors shows superior performance compared to the saliency-driven model and a random gaze selection algorithm. Our results demonstrate that semantic information, from the top-down, influences visual exploration and search tasks significantly, suggesting a potential area of research for integrating it with traditional bottom-up cues.

📄 PDF Abstract BibTeX arXiv:2404.10836

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Humanoid Occupancy: Enabling A Generalized Multimodal Occupancy Perception System on Humanoid Robots

2025-07-27 · Wei Cui, Haoyu Wang, Wenkang Qin, Yijie Guo 외 arxiv

Humanoid robot technology is advancing rapidly, with manufacturers introducing diverse heterogeneous visual perception modules tailored to specific scenarios. Among various perception paradigms, occupancy-based represent…

Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration

2025-02-20 · Pengxiang Ding, Jianfei Ma, Xinyang Tong, Binghong Zou 외

This paper addresses the limitations of current humanoid robot control frameworks, which primarily rely on reactive mechanisms and lack autonomous interaction capabilities due to data scarcity. We propose Humanoid-VLA, a…

Data AugmentationHumanoid ControlMotion Generation

Emergent Active Perception and Dexterity of Simulated Humanoids from Visual Reinforcement Learning

2025-05-18 · Zhengyi Luo, Chen Tessler, Toru Lin, Ye Yuan 외

Human behavior is fundamentally shaped by visual perception -- our ability to interact with the world depends on actively gathering relevant information and adapting our movements accordingly. Behaviors like searching fo…

Object

CLEVER: Stream-based Active Learning for Robust Semantic Perception from Human Instructions

2025-07-21 · Jongseok Lee, Timo Birr, Rudolph Triebel, Tamim Asfour arxiv

We propose CLEVER, an active learning system for robust semantic perception with Deep Neural Networks (DNNs). For data arriving in streams, our system seeks human support when encountering failures and adapts DNNs online…

Active Learning

CyboRacket: A Perception-to-Action Framework for Humanoid Racket Sports

2026-03-15 · Peng Ren, Chuan Qi, Haoyang Ge, Qiyuan Su 외 arxiv

Dynamic ball-interaction tasks remain challenging for robots because they require tight perception-action coupling under limited reaction time. This challenge is especially pronounced in humanoid racket sports, where suc…

Trajectory PredictionVisual Tracking