paper-with-me

홈 › Papers

From the Laboratory to Real-World Application: Evaluating Zero-Shot Scene Interpretation on Edge Devices for Mobile Robotics

2025-11-04 · Nicolas Schuler, Lea Dewald, Nick Baldig, Jürgen Graf arxiv

Video Understanding, Scene Interpretation and Commonsense Reasoning are highly challenging tasks enabling the interpretation of visual information, allowing agents to perceive, interact with and make rational decisions in its environment. Large Language Models (LLMs) and Visual Language Models (VLMs) have shown remarkable advancements in these areas in recent years, enabling domain-specific applications as well as zero-shot open vocabulary tasks, combining multiple domains. However, the required computational complexity poses challenges for their application on edge devices and in the context of Mobile Robotics, especially considering the trade-off between accuracy and inference time. In this paper, we investigate the capabilities of state-of-the-art VLMs for the task of Scene Interpretation and Action Recognition, with special regard to small VLMs capable of being deployed to edge devices in the context of Mobile Robotics. The proposed pipeline is evaluated on a diverse dataset consisting of various real-world cityscape, on-campus and indoor scenarios. The experimental evaluation discusses the potential of these small models on edge devices, with particular emphasis on challenges, weaknesses, inherent model biases and the application of the gained information. Supplementary material is provided via the following repository: https://datahub.rz.rptu.de/hstr-csrl-public/publications/scene-interpretation-on-edge-devices/

📄 PDF Abstract BibTeX arXiv:2511.02427

Code (0)

등록된 구현이 없습니다.

Tasks

Action Recognition

Similar Papers 제목 키워드 기반

Multimodal Rapport Estimation in Real-World HRI

2026-08-19 · Akihiro Sakuramoto, Takato Hayashi, Ryo Miyoshi, Yuki Okafuji 외 arxiv

Evaluating interaction quality in real-world HRI is an important challenge. If interaction quality can be estimated reliably, the results can be used to improve dialogue strategies and ultimately enable robots to adapt t…

Generative Adversarial Networks for Electronic Health Records: A Framework for Exploring and Evaluating Methods for Predicting Drug-Induced Laboratory Test Trajectories

2017-12-01 · Alexandre Yahi, Rami Vanguri, Noémie Elhadad, Nicholas P. Tatonetti

Generative Adversarial Networks (GANs) represent a promising class of generative networks that combine neural networks with game theory. From generating realistic images and videos to assisting musical creation, GANs are…

Predicting Drug-Induced Laboratory Test EffectsRepresentation LearningTime SeriesTime Series Analysis

Computational emotion analysis with multimodal LLMs: Current evidence on an emerging methodological opportunity

2025-12-11 · Hauke Licht arxiv

Research increasingly leverages audio-visual materials to analyze emotions in political communication. Multimodal large language models (mLLMs) promise to enable such analyses through in-context learning. However, we lac…

Associations between depression symptom severity and daily-life gait characteristics derived from long-term acceleration signals in real-world settings

2022-01-29 · Yuezhou Zhang, Amos A Folarin, Shaoxiong Sun, NIcholas Cummins 외

Gait is an essential manifestation of depression. Laboratory gait characteristics have been found to be closely associated with depression. However, the gait characteristics of daily walking in real-world scenarios and t…

LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs

2024-10-18 · Yujun Zhou, Jingdong Yang, Kehan Guo, Pin-Yu Chen 외

Laboratory accidents pose significant risks to human life and property, underscoring the importance of robust safety protocols. Despite advancements in safety training, laboratory personnel may still unknowingly engage i…

BenchmarkingFairnessMultiple-choice