Papers Embodied Question Answering
“Embodied Question Answering” 태그가 달린 논문 40편 · 필터 해제
Enter the Mind Palace: Reasoning and Planning for Long-term Active Embodied Question Answering
As robots become increasingly capable of operating over extended periods -- spanning days, weeks, and even months -- they are expected to accumulate knowledge of their environments and leverage this experience to assist …
Embodied Question AnsweringQuestion AnsweringToSA: Token Merging with Spatial Awareness
Token merging has emerged as an effective strategy to accelerate Vision Transformers (ViT) by reducing computational costs. However, existing methods primarily rely on the visual token's feature similarity for token merg…
Embodied Question AnsweringQuestion AnsweringGeneral-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting
Developing general-purpose navigation policies for unknown environments remains a core challenge in robotics. Most existing systems rely on task-specific neural networks and fixed data flows, limiting generalizability. L…
Embodied Question AnsweringQuestion AnsweringEQA-RM: A Generative Embodied Reward Model with Test-time Scaling
Reward Models (RMs), vital for large model alignment, are underexplored for complex embodied tasks like Embodied Question Answering (EQA) where nuanced evaluation of agents' spatial, temporal, and logical understanding i…
Embodied Question AnsweringQuestion AnsweringMemory-Centric Embodied Question Answer
Embodied Question Answering (EQA) requires agents to autonomously explore and understand the environment to answer context-dependent questions. Existing frameworks typically center around the planner, which guides the st…
Embodied Question AnsweringLarge Language ModelQuestion AnsweringBeyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question Answering
Embodied Question Answering (EQA) is a challenging task in embodied intelligence that requires agents to dynamically explore 3D environments, actively gather visual information, and perform multi-step reasoning to answer…
Embodied Question AnsweringQuestion AnsweringVector Quantized Feature Fields for Fast 3D Semantic Lifting
We generalize lifting to semantic lifting by incorporating per-view masks that indicate relevant pixels for lifting tasks. These masks are determined by querying corresponding multiscale pixel-aligned feature maps, which…
Embodied Question AnsweringQuestion AnsweringCityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space
Embodied Question Answering (EQA) has primarily focused on indoor environments, leaving the complexities of urban settings - spanning environment, action, and perception - largely unexplored. To bridge this gap, we intro…
Embodied Question AnsweringQuestion AnsweringSpatial ReasoningVisual ReasoningTarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
We introduce Tarsier2, a state-of-the-art large vision-language model (LVLM) designed for generating detailed and accurate video descriptions, while also exhibiting superior general video understanding capabilities. Tars…
Embodied Question AnsweringHallucinationLanguage ModelingLanguage Modelling+5GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering
In Embodied Question Answering (EQA), agents must explore and develop a semantic understanding of an unseen environment in order to answer a situated question with confidence. This remains a challenging problem in roboti…
Efficient ExplorationEmbodied Question AnsweringQuestion AnsweringWorld KnowledgeNoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries
The rapid advancement of Vision-Language Models (VLMs) has significantly advanced the development of Embodied Question Answering (EQA), enhancing agents' abilities in language understanding and reasoning within complex a…
BenchmarkingEmbodied Question AnsweringHallucinationQuestion AnsweringTANGO: Training-free Embodied AI Agents for Open-world Tasks
Large Language Models (LLMs) have demonstrated excellent capabilities in composing various modules together to create programs that can perform complex reasoning tasks on images. In this paper, we propose TANGO, an appro…
Embodied Question AnsweringObjectGoal NavigationPointGoal NavigationQuestion AnsweringLSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences
Research on 3D Vision-Language Models (3D-VLMs) is gaining increasing attention, which is crucial for developing embodied AI within 3D scenes, such as visual navigation and embodied question answering. Due to the high de…
Embodied Question AnsweringQuestion AnsweringScene UnderstandingVisual NavigationEfficientEQA: An Efficient Approach for Open Vocabulary Embodied Question Answering
Embodied Question Answering (EQA) is an essential yet challenging task for robotic home assistants. Recent studies have shown that large vision-language models (VLMs) can be effectively utilized for EQA, but existing wor…
Efficient ExplorationEmbodied Question AnsweringQuestion AnsweringRAG+1"Is This It?": Towards Ecologically Valid Benchmarks for Situated Collaboration
We report initial work towards constructing ecologically valid benchmarks to assess the capabilities of large multimodal models for engaging in situated collaboration. In contrast to existing benchmarks, in which questio…
Embodied Question AnsweringQuestion AnsweringvalidMFE-ETP: A Comprehensive Evaluation Benchmark for Multi-modal Foundation Models on Embodied Task Planning
In recent years, Multi-modal Foundation Models (MFMs) and Embodied Artificial Intelligence (EAI) have been advancing side by side at an unprecedented pace. The integration of the two has garnered significant attention fr…
Embodied Question AnsweringQuestion AnsweringTask PlanningMulti-LLM QA with Embodied Exploration
Large language models (LLMs) have grown in popularity due to their natural language interface and pre trained knowledge, leading to rapidly increasing success in question-answering (QA) tasks. More recently, multi-agent …
Embodied Question AnsweringFeature ImportanceQuestion AnsweringMap-based Modular Approach for Zero-shot Embodied Question Answering
Embodied Question Answering (EQA) serves as a benchmark task to evaluate the capability of robots to navigate within novel environments and identify objects in response to human queries. However, existing EQA methods oft…
Embodied Question AnsweringNavigateQuestion AnsweringIs the House Ready For Sleeptime? Generating and Evaluating Situational Queries for Embodied Question Answering
We present and tackle the problem of Embodied Question Answering (EQA) with Situational Queries (S-EQA) in a household environment. Unlike prior EQA work tackling simple queries that directly reference target objects and…
2kEmbodied Question AnsweringHallucinationQuestion Answering+4Explore until Confident: Efficient Exploration for Embodied Question Answering
We consider the problem of Embodied Question Answering (EQA), which refers to settings where an embodied agent such as a robot needs to actively explore an environment to gather information until it is confident about th…
Conformal PredictionEfficient ExplorationEmbodied Question AnsweringQuestion Answering+1