Papers Embodied Question Answering
“Embodied Question Answering” 태그가 달린 논문 40편 · 필터 해제
MEIA: Multimodal Embodied Perception and Interaction in Unknown Environments
With the surge in the development of large language models, embodied intelligence has attracted increasing attention. Nevertheless, prior works on embodied intelligence typically encode scene or historical memory in an u…
Embodied Question AnsweringLanguage ModelingLanguage ModellingLarge Language Model+2OpenEQA: Embodied Question Answering in the Era of Foundation Models
We present a modern formulation of Embodied Question Answering (EQA) as the task of understanding an environment well enough to answer questions about it in natural language. An agent can achieve such an understandin…
Embodied Question AnsweringQuestion AnsweringTowards Learning a Generalist Model for Embodied Navigation
Building a generalist agent that can interact with the world is the intriguing target of AI systems, thus spurring the research for embodied navigation, where an agent is required to navigate according to instructions or…
3D Question Answering (3D-QA)Embodied Question AnsweringNavigateQuestion Answering+1Synthesizing Event-centric Knowledge Graphs of Daily Activities Using Virtual Space
Artificial intelligence (AI) is expected to be embodied in software agents, robots, and cyber-physical systems that can understand the various contextual information of daily life in the home environment to support human…
Decision MakingEmbodied Question AnsweringKnowledge GraphsQuestion AnsweringLLM as A Robotic Brain: Unifying Egocentric Memory and Control
Embodied AI focuses on the study and development of intelligent systems that possess a physical or virtual embodiment (i.e. robots) and are able to dynamically interact with their environment. Memory and control are the …
Embodied Question AnsweringLanguage ModelingLanguage ModellingQuestion Answering+1Explore before Moving: A Feasible Path Estimation and Memory Recalling Framework for Embodied Navigation
An embodied task such as embodied question answering (EmbodiedQA), requires an agent to explore the environment and collect clues to answer a given question that related with specific objects in the scene. The solution o…
Common Sense ReasoningEmbodied Question AnsweringQuestion AnsweringVisual Question Answering (VQA)A Survey of Embodied AI: From Simulators to Research Tasks
There has been an emerging paradigm shift from the era of "internet AI" to "embodied AI", where AI algorithms and agents no longer learn from datasets of images, videos or text curated primarily from the internet. Instea…
Embodied Question AnsweringQuestion AnsweringSurveyVisual NavigationCounterfactual Vision-and-Language Navigation: Unravelling the Unseen
The task of vision-and-language navigation (VLN) requires an agent to follow text instructions to find its way through simulated household environments. A prominent challenge is to train an agent capable of generalising …
counterfactualEmbodied Question AnsweringQuestion AnsweringVision and Language NavigationAllenAct: A Framework for Embodied AI Research
The domain of Embodied AI, in which agents learn to complete tasks through interaction with their environment from egocentric observations, has experienced substantial growth with the advent of deep reinforcement learnin…
Deep Reinforcement LearningEmbodied Question AnsweringInstruction FollowingQuestion AnsweringMulti-Agent Embodied Question Answering in Interactive Environments
We investigate a new AI task --- Multi-Agent Interactive Question Answering --- where several agents explore the scene jointly in interactive environments to answer a question. To cooperate efficiently and answer accurat…
3D ReconstructionEmbodied Question AnsweringQuestion AnsweringSegEQA: Video Segmentation Based Visual Attention for Embodied Question Answering
Embodied Question Answering (EQA) is a newly defined research area where an agent is required to answer the user's questions by exploring the real world environment. It has attracted increasing research interests due to …
Embodied Question AnsweringQuestion AnsweringSegmentationVideo Segmentation+3VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering
Embodied Question Answering (EQA) is a recently proposed task, where an agent is placed in a rich 3D environment and must act based solely on its egocentric input to answer a given question. The desired outcome is that t…
Embodied Question AnsweringQuestion AnsweringReinforcement LearningScene Understanding+1Cross-Task Knowledge Transfer for Visually-Grounded Navigation
Recent efforts on training visual navigation agents conditioned on language using deep reinforcement learning have been successful in learning policies for two different tasks: learning to follow navigational instruction…
Deep Reinforcement LearningDisentanglementEmbodied Question AnsweringQuestion Answering+3Multi-Target Embodied Question Answering
Embodied Question Answering (EQA) is a relatively new task where an agent is asked to answer questions about its environment from egocentric perception. EQA makes the fundamental assumption that every question, e.g., "wh…
Embodied Question AnsweringNavigateQuestion AnsweringVisual Question Answering (VQA)Revisiting EmbodiedQA: A Simple Baseline and Beyond
In Embodied Question Answering (EmbodiedQA), an agent interacts with an environment to gather necessary information for answering user questions. Existing works have laid a solid foundation towards solving this interesti…
Embodied Question AnsweringQuestion AnsweringEmbodied Question Answering in Photorealistic Environments with Point Cloud Perception
To help bridge the gap between internet vision-style problems and the goal of vision for embodied perception we instantiate a large-scale navigation task -- Embodied Question Answering [1] in photo-realistic environments…
Embodied Question AnsweringQuestion AnsweringEmbodied Multimodal Multitask Learning
Recent efforts on training visual navigation agents conditioned on language using deep reinforcement learning have been successful in learning policies for different multimodal tasks, such as semantic goal navigation and…
Deep Reinforcement LearningDisentanglementEmbodied Question AnsweringQuestion Answering+3Blindfold Baselines for Embodied QA
We explore blindfold (question-only) baselines for Embodied Question Answering. The EmbodiedQA task requires an agent to answer a question by intelligently navigating in a simulated environment, gathering necessary visua…
Embodied Question AnsweringQuestion AnsweringNeural Modular Control for Embodied Question Answering
We present a modular approach for learning policies for navigation over long planning horizons from language input. Our hierarchical policy operates at multiple timescales, where the higher-level master policy proposes s…
Embodied Question AnsweringImitation LearningQuestion Answeringreinforcement-learning+2Embodied Question Answering
We present a new AI task -- Embodied Question Answering (EmbodiedQA) -- where an agent is spawned at a random location in a 3D environment and asked a question ("What color is the car?"). In order to answer, the agent mu…
Embodied Question AnsweringNavigateQuestion Answeringreinforcement-learning+2