paper-with-me

홈 › Papers

AI2-THOR: An Interactive 3D Environment for Visual AI

2017-12-14 · Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Matt Deitke, Kiana Ehsani, Daniel Gordon, Yuke Zhu, Aniruddha Kembhavi, Abhinav Gupta, Ali Farhadi

We introduce The House Of inteRactions (THOR), a framework for visual AI research, available at http://ai2thor.allenai.org. AI2-THOR consists of near photo-realistic 3D indoor scenes, where AI agents can navigate in the scenes and interact with objects to perform tasks. AI2-THOR enables research in many different domains including but not limited to deep reinforcement learning, imitation learning, learning by interaction, planning, visual question answering, unsupervised representation learning, object detection and segmentation, and learning models of cognition. The goal of AI2-THOR is to facilitate building visually intelligent models and push the research forward in this domain.

📄 PDF Abstract BibTeX arXiv:1712.05474

Code (2)

allenai/ai2thor 공식 구현
facebookresearch/EgoTV pytorch

Tasks

Deep Reinforcement LearningImitation LearningNavigateobject-detectionObject DetectionQuestion Answeringreinforcement-learningReinforcement LearningReinforcement Learning (RL)Representation LearningVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

RoboTHOR: An Open Simulation-to-Real Embodied AI Platform

2020-04-14 · CVPR 2020 6 · Matt Deitke, Winson Han, Alvaro Herrasti, Aniruddha Kembhavi 외

Visual recognition ecosystems (e.g. ImageNet, Pascal, COCO) have undeniably played a prevailing role in the evolution of modern computer vision. We argue that interactive and embodied visual AI has reached a stage of dev…

IQA: Visual Question Answering in Interactive Environments

2017-12-09 · CVPR 2018 6 · Daniel Gordon, Aniruddha Kembhavi, Mohammad Rastegari, Joseph Redmon 외

We introduce Interactive Question Answering (IQA), the task of answering questions that require an autonomous agent to interact with a dynamic visual environment. IQA presents the agent with a scene and a question, like:…

NavigateReinforcement LearningVisual Question AnsweringVisual Question Answering (VQA)

An Interactive Navigation Method with Effect-oriented Affordance

2024-01-01 · CVPR 2024 1 · Xiaohan Wang, Yuehu Liu, Xinhang Song, Yuyi Liu 외

Visual navigation is to let the agent reach the target according to the continuous visual input. In most previous works visual navigation is usually assumed to be done in a static and ideal environment: the target is…

NavigateVisual Navigation

Pushing it out of the Way: Interactive Visual Navigation

2021-04-28 · CVPR 2021 1 · Kuo-Hao Zeng, Luca Weihs, Ali Farhadi, Roozbeh Mottaghi

We have observed significant progress in visual navigation for embodied agents. A common assumption in studying visual navigation is that the environments are static; this is a limiting assumption. Intelligent navigation…

NavigateVisual Navigation

Orange Lab: Lowering Barriers to Data Mining through Embedded Interactive Workflows

2026-06-08 · Matej Bevec, Aleš Erjavec, Vesna Tanko, Lena Trnovec 외 arxiv

While visual programming of data analysis workflows has become an important vehicle for the democratization of data science, such systems remain largely confined to standalone applications and offer limited support for t…