paper-with-me

Papers

Semantic Audio-Visual Navigation

2020-12-21 · CVPR 2021 1 · Changan Chen, Ziad Al-Halah, Kristen Grauman

Recent work on audio-visual navigation assumes a constantly-sounding target and restricts the role of audio to signaling the target's position. We introduce semantic audio-visual navigation, where objects in the environment make sounds consistent with their semantic meaning (e.g., toilet flushing, door creaking) and acoustic events are sporadic or short in duration. We propose a transformer-based model to tackle this new semantic AudioGoal task, incorporating an inferred goal descriptor that captures both spatial and semantic properties of the target. Our model's persistent multimodal memory enables it to reach the goal even long after the acoustic event stops. In support of the new task, we also expand the SoundSpaces audio simulations to provide semantically grounded sounds for an array of objects in Matterport3D. Our method strongly outperforms existing audio-visual navigation methods by learning to associate semantic, acoustic, and visual cues.

📄 PDF Abstract BibTeX arXiv:2012.11583

Code (0)

등록된 구현이 없습니다.

Tasks

PositionVisual Navigation

Similar Papers 제목 키워드 기반

Knowledge-driven Scene Priors for Semantic Audio-Visual Embodied Navigation

2022-12-21 · Gyan Tatiya, Jonathan Francis, Luca Bondi, Ingrid Navarro 외

Generalisation to unseen contexts remains a challenge for embodied navigation agents. In the context of semantic audio-visual navigation (SAVi) tasks, the notion of generalisation should include both generalising to unse…

Visual Navigation

Pay Self-Attention to Audio-Visual Navigation

2022-10-04 · Yinfeng Yu, Lele Cao, Fuchun Sun, Xiaohong Liu 외

Audio-visual embodied navigation, as a hot research topic, aims training a robot to reach an audio target using egocentric visual (from the sensors mounted on the robot) and audio (emitted from the target) input. The aud…

Visual Navigation

AVLEN: Audio-Visual-Language Embodied Navigation in 3D Environments

2022-10-14 · Sudipta Paul, Amit K. Roy-Chowdhury, Anoop Cherian

Recent years have seen embodied visual navigation advance in two distinct directions: (i) in equipping the AI agent to follow natural language instructions, and (ii) in making the navigable world multimodal, e.g., audio-…

AI AgentHierarchical Reinforcement LearningNavigateVisual Navigation

Semantic Audio-Visual Navigation in Continuous Environments

2026-03-20 · Yichen Zeng, Hebaixu Wang, Meng Liu, Yu Zhou 외 arxiv

Audio-visual navigation enables embodied agents to navigate toward sound-emitting targets by leveraging both auditory and visual cues. However, most existing approaches rely on precomputed room impulse responses (RIRs) f…

Visual Navigation

LH-AVLN: A Benchmark for Long-Horizon Audio-Visual-Language Navigation

2026-07-04 · Rufeng Chen, Yue Chang, Zili Shao, Zhaofan Zhang 외 arxiv

Embodied navigation is moving toward long-horizon missions, yet existing long-horizon benchmarks are largely acoustically silent, and audio-visual navigation tasks typically focus on a single goal. We introduce LH-AVLN, …

Visual Navigation