paper-with-me

Papers

Bridging Visual Perception with Contextual Semantics for Understanding Robot Manipulation Tasks

2019-09-16 · Chen Jiang, Martin Jagersand

Understanding manipulation scenarios allows intelligent robots to plan for appropriate actions to complete a manipulation task successfully. It is essential for intelligent robots to semantically interpret manipulation knowledge by describing entities, relations and attributes in a structural manner. In this paper, we propose an implementing framework to generate high-level conceptual dynamic knowledge graphs from video clips. A combination of a Vision-Language model and an ontology system, in correspondence with visual perception and contextual semantics, is used to represent robot manipulation knowledge with Entity-Relation-Entity (E-R-E) and Entity-Attribute-Value (E-A-V) tuples. The proposed method is flexible and well-versed. Using the framework, we present a case study where robot performs manipulation actions in a kitchen environment, bridging visual perception with contextual semantics using the generated dynamic knowledge graphs.

📄 PDF Abstract BibTeX arXiv:1909.07459

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeCommon Sense ReasoningKnowledge GraphsLanguage ModelingLanguage ModellingRobot Manipulation

Similar Papers 제목 키워드 기반

LOIS: Looking Out of Instance Semantics for Visual Question Answering

2023-07-26 · Siyu Zhang, Yeming Chen, Yaoru Sun, Fang Wang 외

Visual question answering (VQA) has been intensively studied as a multimodal task that requires effort in bridging vision and language to infer answers correctly. Recent attempts have developed various attention-based mo…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Visual Reasoning

The Roles of Contextual Semantic Relevance Metrics in Human Visual Processing

2024-10-13 · Kun Sun, Rong Wang

Semantic relevance metrics can capture both the inherent semantics of individual objects and their relationships to other elements within a visual scene. Numerous previous research has demonstrated that these metrics can…

VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs

2025-09-30 · Peng Liu, Haozhan Shen, Chunxin Fang, Zhicheng Sun 외 arxiv

Vision-Language Models (VLMs) excel at high-level scene understanding but falter on fine-grained perception tasks requiring precise localization. This failure stems from a fundamental mismatch, as generating exact numeri…

Scene UnderstandingVisual Grounding

RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding

2025-11-27 · Xiyan Liu, Han Wang, Yuhu Wang, Junjie Cai 외 arxiv

Understanding mid-level road semantics, which capture the structural and contextual cues that link low-level perception to high-level planning, is essential for reliable autonomous driving and digital map construction. H…

Scene UnderstandingAutonomous DrivingVisual Reasoning

SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models

2026-03-17 · Tianyu Xie, Jinfa Huang, Yuexiao Ma, Rongfang Luo 외 arxiv

Omni-modal large language models (OLMs) redefine human-machine interaction by natively integrating audio, vision, and text. However, existing OLM benchmarks remain anchored to static, accuracy-centric tasks, leaving a cr…