Bridging Visual Perception with Contextual Semantics for Understanding Robot Manipulation Tasks
Understanding manipulation scenarios allows intelligent robots to plan for appropriate actions to complete a manipulation task successfully. It is essential for intelligent robots to semantically interpret manipulation knowledge by describing entities, relations and attributes in a structural manner. In this paper, we propose an implementing framework to generate high-level conceptual dynamic knowledge graphs from video clips. A combination of a Vision-Language model and an ontology system, in correspondence with visual perception and contextual semantics, is used to represent robot manipulation knowledge with Entity-Relation-Entity (E-R-E) and Entity-Attribute-Value (E-A-V) tuples. The proposed method is flexible and well-versed. Using the framework, we present a case study where robot performs manipulation actions in a kitchen environment, bridging visual perception with contextual semantics using the generated dynamic knowledge graphs.
Code (0)
등록된 구현이 없습니다.
Tasks
AttributeCommon Sense ReasoningKnowledge GraphsLanguage ModelingLanguage ModellingRobot ManipulationSimilar Papers 제목 키워드 기반
LOIS: Looking Out of Instance Semantics for Visual Question Answering
Visual question answering (VQA) has been intensively studied as a multimodal task that requires effort in bridging vision and language to infer answers correctly. Recent attempts have developed various attention-based mo…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Visual ReasoningThe Roles of Contextual Semantic Relevance Metrics in Human Visual Processing
Semantic relevance metrics can capture both the inherent semantics of individual objects and their relationships to other elements within a visual scene. Numerous previous research has demonstrated that these metrics can…
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
Vision-Language Models (VLMs) excel at high-level scene understanding but falter on fine-grained perception tasks requiring precise localization. This failure stems from a fundamental mismatch, as generating exact numeri…
Scene UnderstandingVisual GroundingRoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding
Understanding mid-level road semantics, which capture the structural and contextual cues that link low-level perception to high-level planning, is essential for reliable autonomous driving and digital map construction. H…
Scene UnderstandingAutonomous DrivingVisual ReasoningSocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
Omni-modal large language models (OLMs) redefine human-machine interaction by natively integrating audio, vision, and text. However, existing OLM benchmarks remain anchored to static, accuracy-centric tasks, leaving a cr…