paper-with-me

Papers

A large scale multi-view RGBD visual affordance learning dataset

2022-03-26 · Zeyad Khalifa, Syed Afaq Ali Shah

The physical and textural attributes of objects have been widely studied for recognition, detection and segmentation tasks in computer vision.~A number of datasets, such as large scale ImageNet, have been proposed for feature learning using data hungry deep neural networks and for hand-crafted feature extraction. To intelligently interact with objects, robots and intelligent machines need the ability to infer beyond the traditional physical/textural attributes, and understand/learn visual cues, called visual affordances, for affordance recognition, detection and segmentation. To date there is no publicly available large dataset for visual affordance understanding and learning. In this paper, we introduce a large scale multi-view RGBD visual affordance learning dataset, a benchmark of 47210 RGBD images from 37 object categories, annotated with 15 visual affordance categories. To the best of our knowledge, this is the first ever and the largest multi-view RGBD visual affordance learning dataset. We benchmark the proposed dataset for affordance segmentation and recognition tasks using popular Vision Transformer and Convolutional Neural Networks. Several state-of-the-art deep learning networks are evaluated each for affordance recognition and segmentation tasks. Our experimental results showcase the challenging nature of the dataset and present definite prospects for new and robust affordance learning algorithms. The dataset is publicly available at https://sites.google.com/view/afaqshah/dataset.

📄 PDF Abstract BibTeX arXiv:2203.14092

Code (0)

등록된 구현이 없습니다.

Tasks

Affordance RecognitionSegmentation

Similar Papers 제목 키워드 기반

Resource-Efficient RGBD Aerial Tracking

2023-01-01 · CVPR 2023 1 · Jinyu Yang, Shang Gao, Zhe Li, Feng Zheng 외

Aerial robots are now able to fly in complex environments, and drone-captured data gains lots of attention in object tracking. However, current research on aerial perception has mainly focused on limited categories, …

GPUObject Tracking

Refer-it-in-RGBD: A Bottom-up Approach for 3D Visual Grounding in RGBD Images

2021-03-14 · CVPR 2021 1 · Haolin Liu, Anran Lin, Xiaoguang Han, Lei Yang 외

Grounding referring expressions in RGBD image has been an emerging field. We present a novel task of 3D visual grounding in single-view RGBD image where the referred objects are often only partially scanned due to occlus…

3D visual groundingObjectVisual Grounding

VTGaussian-SLAM: RGBD SLAM for Large Scale Scenes with Splatting View-Tied 3D Gaussians

2025-06-03 · Pengchong Hu, Zhizhong Han

Jointly estimating camera poses and mapping scenes from RGBD images is a fundamental task in simultaneous localization and mapping (SLAM). State-of-the-art methods employ 3D Gaussians to represent a scene, and render the…

GPUSimultaneous Localization and Mapping

RGBD1K: A Large-scale Dataset and Benchmark for RGB-D Object Tracking

2022-08-21 · Xue-Feng Zhu, Tianyang Xu, Zhangyong Tang, Zucheng Wu 외

RGB-D object tracking has attracted considerable attention recently, achieving promising performance thanks to the symbiosis between visual and depth channels. However, given a limited amount of annotated RGB-D tracking …

Object TrackingVisual Object Tracking

RGBD2: Generative Scene Synthesis via Incremental View Inpainting using RGBD Diffusion Models

2022-12-12 · CVPR 2023 1 · Jiabao Lei, Jiapeng Tang, Kui Jia

We address the challenge of recovering an underlying scene geometry and colors from a sparse set of RGBD view observations. In this work, we present a new solution termed RGBD$^2$ that sequentially generates novel RGBD v…