paper-with-me

Papers

BARISTA: A Multi-Task Egocentric Benchmark for Compositional Visual Understanding

2026-05-12 · Patrick Knab, Orgest Xhelili, Inis Buzi, Drago Andres Guggiana Nilo, Mohd Saquib Khan, Lorenz Kolb, Manuel Scherzer, Kerem Yildirir, Christian Bartelt, Philipp Johannes Schubert arxiv

Scene understanding is central to general physical intelligence, and video is a primary modality for capturing both state and temporal dynamics of a scene. Yet understanding physical processes remains difficult, as models must combine object localization, hand-object interactions, relational parsing, temporal reasoning, and step-level procedural inference. Existing benchmarks usually evaluate these capabilities separately, limiting diagnosis of why models fail on procedural tasks. We introduce BARISTA, a densely annotated egocentric dataset and benchmark of 185 real-world coffee-preparation videos covering fully automatic, portafilter-based, and capsule-based workflows. BARISTA provides verified per-frame scene graphs linking persistent object identities to masks, tracks, boxes, attributes, typed relations, hand-object interactions, activities, and process steps. From these graphs, we derive zero-shot language-based tasks spanning phrase grounding, hand-object interaction recognition, referring, activity recognition, relation extraction, and temporal visual question answering. Experiments reveal strong variation across task families and no consistently dominant model family, positioning BARISTA as a challenging diagnostic benchmark for procedural video understanding. Code and dataset available at https://huggingface.co/datasets/ramblr/BARISTA.

📄 PDF Abstract BibTeX arXiv:2605.12074

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringActivity RecognitionObject LocalizationScene Understanding

Similar Papers 제목 키워드 기반

EgoTV: Egocentric Task Verification from Natural Language Task Descriptions

2023-03-29 · ICCV 2023 1 · Rishi Hazra, Brian Chen, Akshara Rai, Nitin Kamra 외

To enable progress towards egocentric agents capable of understanding everyday tasks specified in natural language, we propose a benchmark and a synthetic dataset called Egocentric Task Verification (EgoTV). The goal in …

COMBO: Compositional World Models for Embodied Multi-Agent Cooperation

2024-04-16 · Hongxin Zhang, Zeyuan Wang, Qiushi Lyu, Zheyuan Zhang 외

In this paper, we investigate the problem of embodied multi-agent cooperation, where decentralized agents must cooperate given only egocentric views of the world. To effectively plan in this setting, in contrast to learn…

Barista - a Graphical Tool for Designing and Training Deep Neural Networks

2018-02-13 · Soeren Klemm, Aaron Scherzinger, Dominik Drees, Xiaoyi Jiang

In recent years, the importance of deep learning has significantly increased in pattern recognition, computer vision, and artificial intelligence research, as well as in industry. However, despite the existence of multip…

Deep Learning

Alquist 5.0: Dialogue Trees Meet Generative Models. A Novel Approach for Enhancing SocialBot Conversations

2023-10-24 · Ondřej Kobza, Jan Čuhel, Tommaso Gargiani, David Herel 외

We present our SocialBot -- Alquist~5.0 -- developed for the Alexa Prize SocialBot Grand Challenge~5. Building upon previous versions of our system, we introduce the NRG Barista and outline several innovative approaches …

Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images

2026-05-12 · Yuangong Chen, Wai Keung Wong, Jiaxing Li, Ioannis Patras 외 arxiv

Multimodal Large Language Models (MLLMs) show strong visual perception, yet remain limited in reasoning about space under changing viewpoints. We study this challenge as Perspective-Conditioned Spatial Reasoning (PCSR) i…

Spatial ReasoningObject Counting