VoxML: A Visualization Modeling Language
We present the specification for a modeling language, VoxML, which encodes semantic knowledge of real-world objects represented as three-dimensional models, and of events and attributes related to and enacted over these objects. VoxML is intended to overcome the limitations of existing 3D visual markup languages by allowing for the encoding of a broad range of semantic knowledge that can be exploited by a variety of systems and platforms, leading to multimodal simulations of real-world scenarios using conceptual objects that represent their semantic values.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
An Abstract Specification of VoxML as an Annotation Language
VoxML is a modeling language used to map natural language expressions into real-time visualizations using commonsense semantic knowledge of objects and events. Its utility has been demonstrated in embodied simulation env…
Human Agent CollaborationHuman-Object Interaction DetectionObjectBuilding Multimodal Simulations for Natural Language
In this tutorial, we introduce a computational framework and modeling language (VoxML) for composing multimodal simulations of natural language expressions within a 3D simulation environment (VoxSim). We demonstrate how …
Formal LogicReferring ExpressionReferring expression generationScene Generation+1ECAT: Event Capture Annotation Tool
This paper introduces the Event Capture Annotation Tool (ECAT), a user-friendly, open-source interface tool for annotating events and their participants in video, capable of extracting the 3D positions and orientations o…
AttributeThe VoxWorld Platform for Multimodal Embodied Agents
We present a five-year retrospective on the development of the VoxWorld platform, first introduced as a multimodal platform for modeling motion language, that has evolved into a platform for rapidly building and deployin…
multimodal interactionThe Development of Multimodal Lexical Resources
Human communication is a multimodal activity, involving not only speech and written expressions, but intonation, images, gestures, visual clues, and the interpretation of actions through perception. In this paper, we des…
Question AnsweringVisual Question Answering (VQA)