paper-with-me

Papers

Multimodal Large Language Models for Real-Time Situated Reasoning

2026-02-02 · Giulio Antonio Abbo, Senne Lenaerts, Tony Belpaeme arxiv

In this work, we explore how multimodal large language models can support real-time context- and value-aware decision-making. To do so, we combine the GPT-4o language model with a TurtleBot 4 platform simulating a smart vacuum cleaning robot in a home. The model evaluates the environment through vision input and determines whether it is appropriate to initiate cleaning. The system highlights the ability of these models to reason about domestic activities, social norms, and user preferences and take nuanced decisions aligned with the values of the people involved, such as cleanliness, comfort, and safety. We demonstrate the system in a realistic home environment, showing its ability to infer context and values from limited visual input. Our results highlight the promise of multimodal large language models in enhancing robotic autonomy and situational awareness, while also underscoring challenges related to consistency, bias, and real-time performance.

📄 PDF Abstract BibTeX arXiv:2602.01880

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Modern System Recipe for Situated Embodied Human-Robot Conversation with Real-Time Multimodal LLMs and Tool-Calling

2026-02-04 · Dong Won Lee, Sarah Gillet, Louis-Philippe Morency, Cynthia Breazeal 외 arxiv

Situated embodied conversation requires robots to interleave real-time dialogue with active perception: deciding what to look at, when to look, and what to say under tight latency constraints. We present a simple, minima…

TRACE: Real-Time Multimodal Common Ground Tracking in Situated Collaborative Dialogues

2025-03-12 · Hannah VanderHoeven, Brady Bhalla, Ibrahim Khebour, Austin Youngren 외

We present TRACE, a novel system for live *common ground* tracking in situated collaborative tasks. With a focus on fast, real-time performance, TRACE tracks the speech, actions, gestures, and visual attention of partici…

Position

"Is This It?": Towards Ecologically Valid Benchmarks for Situated Collaboration

2024-08-30 · Dan Bohus, Sean Andrist, Yuwei Bao, Eric Horvitz 외

We report initial work towards constructing ecologically valid benchmarks to assess the capabilities of large multimodal models for engaging in situated collaboration. In contrast to existing benchmarks, in which questio…

Embodied Question AnsweringQuestion Answeringvalid

UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

2023-10-08 · Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye 외

Text is ubiquitous in our visual world, conveying crucial information, such as in documents, websites, and everyday photographs. In this work, we propose UReader, a first exploration of universal OCR-free visually-situat…

DecoderLanguage ModelingLanguage ModellingLarge Language Model+2

SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking

2025-05-25 · Junnan Liu, Linhao Luo, Thuy-Trang Vu, Gholamreza Haffari

Recent advances in large language models (LLMs) demonstrate their impressive reasoning capabilities. However, the reasoning confined to internal parametric space limits LLMs' access to real-time information and understan…

Mathematical ReasoningMulti-hop Question AnsweringQuestion Answeringtext-based games