paper-with-me

Papers

THDA: Treasure Hunt Data Augmentation for Semantic Navigation

2021-01-01 · ICCV 2021 10 · Oleksandr Maksymets, Vincent Cartillier, Aaron Gokaslan, Erik Wijmans, Wojciech Galuba, Stefan Lee, Dhruv Batra

Can general-purpose neural models learn to navigate? For PointGoal navigation (""go to x, y""), the answer is a clear `yes' -- mapless neural models composed of task-agnostic components (CNNs and RNNs) trained with large-scale model-free reinforcement learning achieve near-perfect performance. However, for ObjectGoal navigation (""find a TV""), this is an open question; one we tackle in this paper. The current best-known result on ObjectNav with general-purpose models is 6% success rate. First, we show that the key problem is overfitting. Large-scale training results in 94% success rate on training environments and only 8% in validation. We observe that this stems from agents memorizing environment layouts during training -- sidestepping the need for exploration and directly learning shortest paths to nearby goal objects. We show that this is a natural consequence of optimizing for the task metric (which in fact penalizes exploration), is enabled by powerful observation encoders, and is possible due to the finite set of training environment configurations. Informed by our findings, we introduce Treasure Hunt Data Augmentation (THDA) to address overfitting in ObjectNav. THDA inserts 3D scans of household objects at arbitrary scene locations and uses them as ObjectNav goals -- augmenting and greatly expanding the set of training layouts. Taken together with our other proposed changes, we improve the state of art on the Habitat ObjectGoal Navigation benchmark by 90% (from 14% success rate to 27%) and path efficiency by 48% (from 7.5 SPL to 11.1 SPL).

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationNavigateObjectGoal NavigationOpen-Ended Question AnsweringPointGoal Navigation

Similar Papers 제목 키워드 기반

Adventurer's Treasure Hunt: A Transparent System for Visually Grounded Compositional Visual Question Answering based on Scene Graphs

2021-06-28 · Daniel Reich, Felix Putze, Tanja Schultz

With the expressed goal of improving system transparency and visual grounding in the reasoning process in VQA, we present a modular system for the task of compositional VQA based on scene graphs. Our system is called "Ad…

Question AnsweringTask 2Visual GroundingVisual Question Answering+1

Hunting for Polluted White Dwarfs and Other Treasures with Gaia XP Spectra and Unsupervised Machine Learning

2024-05-27 · Malia L. Kao, Keith Hawkins, Laura K. Rogers, Amy Bonsor 외

White dwarfs (WDs) polluted by exoplanetary material provide the unprecedented opportunity to directly observe the interiors of exoplanets. However, spectroscopic surveys are often limited by brightness constraints, and …

Diversity

Cheryl's Birthday

2017-07-27 · Hans van Ditmarsch, Michael Ian Hartley, Barteld Kooi, Jonathan Welton 외

We present four logic puzzles and after that their solutions. Joseph Yeo designed 'Cheryl's Birthday'. Mike Hartley came up with a novel solution for 'One Hundred Prisoners and a Light Bulb'. Jonathan Welton designed 'A …

AirHunt: Bridging VLM Semantics and Continuous Planning for Efficient Aerial Object Navigation

2026-01-19 · Xuecheng Chen, Zongzhuo Liu, Jianfa Ma, Bang Du 외 arxiv

Recent advances in large Vision-Language Models (VLMs) have provided rich semantic understanding that empowers drones to search for open-set objects via natural language instructions. However, prior systems struggle to i…

Zero-shot GeneralizationScene Understanding

Dynamic Concept Composition for Zero-Example Event Detection

2016-01-14 · Xiaojun Chang, Yi Yang, Guodong Long, Chengqi Zhang 외

In this paper, we focus on automatically detecting events in unconstrained videos without the use of any visual training exemplars. In principle, zero-shot learning makes it possible to train an event detection model bas…

Event DetectionZero-Shot Learning