paper-with-me

홈 › Papers

Situational Grounding within Multimodal Simulations

2019-02-05 · James Pustejovsky, Nikhil Krishnaswamy

In this paper, we argue that simulation platforms enable a novel type of embodied spatial reasoning, one facilitated by a formal model of object and event semantics that renders the continuous quantitative search space of an open-world, real-time environment tractable. We provide examples for how a semantically-informed AI system can exploit the precise, numerical information provided by a game engine to perform qualitative reasoning about objects and events, facilitate learning novel concepts from data, and communicate with a human to improve its models and demonstrate its understanding. We argue that simulation environments, and game engines in particular, bring together many different notions of "simulation" and many different technologies to provide a highly-effective platform for developing both AI systems and tools to experiment in both machine and human intelligence.

📄 PDF Abstract BibTeX arXiv:1902.01886

Code (0)

등록된 구현이 없습니다.

Tasks

Novel ConceptsSpatial Reasoning

Similar Papers 제목 키워드 기반

RAPTOR-AI for Disaster OODA Loop: Hierarchical Multimodal RAG with Experience-Driven Agentic Decision-Making

2026-01-18 · Takato Yasuno arxiv

Humanitarian Assistance and Disaster Relief (HADR) operations demand rapid synthesis of multimodal information for time-critical decision-making under extreme uncertainty. Traditional information systems struggle with th…

Multimodal Situational Safety

2024-10-08 · Kaiwen Zhou, Chengzhi Liu, Xuandong Zhao, Anderson Compalas 외

Multimodal Large Language Models (MLLMs) are rapidly evolving, demonstrating impressive capabilities as multimodal assistants that interact with both humans and their environments. However, this increased sophistication …

Instruction Following

Using Indirect Encoding of Multiple Brains to Produce Multimodal Behavior

2016-04-26 · Jacob Schrum, Joel Lehman, Sebastian Risi

An important challenge in neuroevolution is to evolve complex neural networks with multiple modes of behavior. Indirect encodings can potentially answer this challenge. Yet in practice, indirect encodings do not yield ef…

GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models

2023-11-15 · Serwan Jassim, Mario Holubar, Annika Richter, Cornelius Wolff 외

This paper presents GRASP, a novel benchmark to evaluate the language grounding and physical understanding capabilities of video-based multimodal large language models (LLMs). This evaluation is accomplished via a two-ti…

Unity

MERGE: Guided Vision-Language Models for Multi-Actor Event Reasoning and Grounding in Human-Robot Interaction

2026-03-19 · Joerg Deigmoeller, Nakul Agarwal, Stephan Hasler, Daniel Tanneberg 외 arxiv

We introduce MERGE, a system for situational grounding of actors, objects, and events in dynamic human-robot group interactions. Effective collaboration in such settings requires consistent situational awareness, built o…

Zero-shot Generalization