Kefa: A Knowledge Enhanced and Fine-grained Aligned Speaker for Navigation Instruction Generation
We introduce a novel speaker model \textsc{Kefa} for navigation instruction generation. The existing speaker models in Vision-and-Language Navigation suffer from the large domain gap of vision features between different environments and insufficient temporal grounding capability. To address the challenges, we propose a Knowledge Refinement Module to enhance the feature representation with external knowledge facts, and an Adaptive Temporal Alignment method to enforce fine-grained alignment between the generated instructions and the observation sequences. Moreover, we propose a new metric SPICE-D for navigation instruction evaluation, which is aware of the correctness of direction phrases. The experimental results on R2R and UrbanWalk datasets show that the proposed KEFA speaker achieves state-of-the-art instruction generation performance for both indoor and outdoor scenes.
Code (1)
Tasks
Vision and Language NavigationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Knowledge-Guided Prompt Learning for Deepfake Facial Image Detection
Recent generative models demonstrate impressive performance on synthesizing photographic images, which makes humans hardly to distinguish them from pristine ones, especially on realistic-looking synthetic facial images. …
Face SwappingPrompt LearningTowards Coarse-grained Visual Language Navigation Task Planning Enhanced by Event Knowledge Graph
Visual language navigation (VLN) is one of the important research in embodied AI. It aims to enable an agent to understand the surrounding environment and complete navigation tasks. VLN instructions could be categorized …
Language ModellingLarge Language ModelTask PlanningDeepPerception: Advancing R1-like Cognitive Visual Perception in MLLMs for Knowledge-Intensive Visual Grounding
Human experts excel at fine-grained visual discrimination by leveraging domain knowledge to refine perceptual features, a capability that remains underdeveloped in current Multimodal Large Language Models (MLLMs). Despit…
Domain GeneralizationMultimodal ReasoningVisual GroundingKLMo: Knowledge Graph Enhanced Pretrained Language Model with Fine-Grained Relationships
Interactions between entities in knowledge graph (KG) provide rich knowledge for language representation learning. However, existing knowledge-enhanced pretrained language models (PLMs) only focus on entity information a…
Entity LinkingEntity TypingLanguage ModelingLanguage Modelling+4FineFake: A Knowledge-Enriched Dataset for Fine-Grained Multi-Domain Fake News Detection
Existing benchmarks for fake news detection have significantly contributed to the advancement of models in assessing the authenticity of news content. However, these benchmarks typically focus solely on news pertaining t…
Domain AdaptationFake News Detection