paper-with-me

Papers

Interacted Object Grounding in Spatio-Temporal Human-Object Interactions

2024-12-27 · Xiaoyang Liu, Boran Wen, Xinpeng Liu, Zizheng Zhou, Hongwei Fan, Cewu Lu, Lizhuang Ma, Yulong Chen, Yong-Lu Li

Spatio-temporal Human-Object Interaction (ST-HOI) understanding aims at detecting HOIs from videos, which is crucial for activity understanding. However, existing whole-body-object interaction video benchmarks overlook the truth that open-world objects are diverse, that is, they usually provide limited and predefined object classes. Therefore, we introduce a new open-world benchmark: Grounding Interacted Objects (GIO) including 1,098 interacted objects class and 290K interacted object boxes annotation. Accordingly, an object grounding task is proposed expecting vision systems to discover interacted objects. Even though today's detectors and grounding methods have succeeded greatly, they perform unsatisfactorily in localizing diverse and rare objects in GIO. This profoundly reveals the limitations of current vision systems and poses a great challenge. Thus, we explore leveraging spatio-temporal cues to address object grounding and propose a 4D question-answering framework (4D-QA) to discover interacted objects from diverse videos. Our method demonstrates significant superiority in extensive experiments compared to current baselines. Data and code will be publicly available at https://github.com/DirtyHarryLYL/HAKE-AVA.

📄 PDF Abstract BibTeX arXiv:2412.19542

Code (1)

dirtyharrylyl/hake-ava 공식 구현 pytorch

Tasks

Human-Object Interaction DetectionObjectQuestion Answering

Similar Papers 제목 키워드 기반

Discovering A Variety of Objects in Spatio-Temporal Human-Object Interactions

2022-11-14 · Yong-Lu Li, Hongwei Fan, Zuoyu Qiu, Yiming Dou 외

Spatio-temporal Human-Object Interaction (ST-HOI) detection aims at detecting HOIs from videos, which is crucial for activity understanding. In daily HOIs, humans often interact with a variety of objects, e.g., holding a…

Human-Object Interaction DetectionObjectobject-detectionObject Detection+1

Spatio-Temporal Interaction Graph Parsing Networks for Human-Object Interaction Recognition

2021-08-19 · Ning Wang, Guangming Zhu, Liang Zhang, Peiyi Shen 외

For a given video-based Human-Object Interaction scene, modeling the spatio-temporal relationship between humans and objects are the important cue to understand the contextual information presented in the video. With the…

Human-Object Interaction DetectionObject

Hand-Object Interaction Reasoning

2022-01-13 · Jian Ma, Dima Damen

This paper proposes an interaction reasoning network for modelling spatio-temporal relationships between hands and objects in video. The proposed interaction unit utilises a Transformer module to reason about each acting…

Action RecognitionObject

Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences

2020-01-19 · CVPR 2020 6 · Zhu Zhang, Zhou Zhao, Yang Zhao, Qi. Wang 외

In this paper, we consider a novel task, Spatio-Temporal Video Grounding for Multi-Form Sentences (STVG). Given an untrimmed video and a declarative/interrogative sentence depicting an object, STVG aims to localize the s…

FormObjectSentenceSpatio-Temporal Video Grounding+1

Skeleton-Based Mutually Assisted Interacted Object Localization and Human Action Recognition

2021-10-28 · Liang Xu, Cuiling Lan, Wenjun Zeng, Cewu Lu

Skeleton data carries valuable motion information and is widely explored in human action recognition. However, not only the motion information but also the interaction with the environment provides discriminative cues to…

Action RecognitionObjectObject LocalizationTemporal Action Localization