paper-with-me

Papers

Sonicverse: A Multisensory Simulation Platform for Embodied Household Agents that See and Hear

2023-06-01 · Ruohan Gao, Hao Li, Gokul Dharan, Zhuzhu Wang, Chengshu Li, Fei Xia, Silvio Savarese, Li Fei-Fei, Jiajun Wu

Developing embodied agents in simulation has been a key research topic in recent years. Exciting new tasks, algorithms, and benchmarks have been developed in various simulators. However, most of them assume deaf agents in silent environments, while we humans perceive the world with multiple senses. We introduce Sonicverse, a multisensory simulation platform with integrated audio-visual simulation for training household agents that can both see and hear. Sonicverse models realistic continuous audio rendering in 3D environments in real-time. Together with a new audio-visual VR interface that allows humans to interact with agents with audio, Sonicverse enables a series of embodied AI tasks that need audio-visual perception. For semantic audio-visual navigation in particular, we also propose a new multi-task learning model that achieves state-of-the-art performance. In addition, we demonstrate Sonicverse's realism via sim-to-real transfer, which has not been achieved by other simulators: an agent trained in Sonicverse can successfully perform audio-visual navigation in real-world environments. Sonicverse is available at: https://github.com/StanfordVL/Sonicverse.

📄 PDF Abstract BibTeX arXiv:2306.00923

Code (1)

stanfordvl/sonicverse 공식 구현

Tasks

Multi-Task LearningVisual Navigation

Similar Papers 제목 키워드 기반

A Scalable Embodied Intelligence Platform for Seamless Real-to-Sim-to-Real Transfer of Household Mobile Manipulation Tasks

2026-06-17 · Kui Yang, Xianlei Long, Haoxuan Li, Yan Ding 외 arxiv

Mobile manipulation is a fundamental capability in embodied intelligence robotics. The growing demand for robust and generalizable manipulation in unstructured household environments has driven rapid progress in embodied…

Scene Generation

MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World

2024-01-16 · CVPR 2024 1 · Yining Hong, Zishuo Zheng, Peihao Chen, Yian Wang 외

Human beings possess the capability to multiply a melange of multisensory cues while actively exploring and interacting with the 3D world. Current multi-modal large language models, however, passively absorb sensory data…

Language ModelingLanguage ModellingLarge Language Model

Neural Multisensory Scene Inference

2019-10-06 · NeurIPS 2019 12 · Jae Hyun Lim, Pedro O. Pinheiro, Negar Rostamzadeh, Christopher Pal 외

For embodied agents to infer representations of the underlying 3D physical world they inhabit, they should efficiently combine multisensory cues from numerous trials, e.g., by looking at and touching objects. Despite its…

Computational EfficiencyRepresentation Learning

The ObjectFolder Benchmark: Multisensory Learning with Neural and Real Objects

2023-06-01 · CVPR 2023 1 · Ruohan Gao, Yiming Dou, Hao Li, Tanmay Agarwal 외

We introduce the ObjectFolder Benchmark, a benchmark suite of 10 tasks for multisensory object-centric learning, centered around object recognition, reconstruction, and manipulation with sight, sound, and touch. We also …

BenchmarkingObjectObject Recognition

MarketGen: A Scalable Simulation Platform with Auto-Generated Embodied Supermarket Environments

2025-11-26 · Xu Hu, Yiyang Feng, Junran Peng, Jiawei He 외 arxiv

The development of embodied agents for complex commercial environments is hindered by a critical gap in existing robotics datasets and benchmarks, which primarily focus on household or tabletop settings with short-horizo…

Scene Generation