Human Interaction Recognition
8개 벤치마크 · 논문 23편 · 이 태스크의 논문 보기 →
Benchmarks
NTU RGB+D 120
NTU RGB+D
UT
BIT
SBU / SBU-Refine
EPIC-SOUNDS
SBU
UT-Interaction
Most implemented
Slow-Fast Auditory Streams For Audio Recognition
CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action Recognition
Empathic Grounding: Explorations using Multimodal Interaction and Large Language Models with Conversational Agents
SkateFormer: Skeletal-Temporal Transformer for Human Action Recognition
Papers
Learning Adaptive Node Selection with External Attention for Human Interaction Recognition
Most GCN-based methods model interacting individuals as independent graphs, neglecting their inherent inter-dependencies. Although recent approaches utilize predefined interaction adjacency matrices to integrate particip…
Human Interaction RecognitionDynamic Scene Understanding from Vision-Language Representations
Images depicting complex, dynamic scenes are challenging to parse automatically, requiring both high-level comprehension of the overall situation and fine-grained identification of participating entities and their intera…
Grounded Situation RecognitionHuman-Human Interaction RecognitionHuman Interaction RecognitionHuman-Object Interaction Detection+2OV-HHIR: Open Vocabulary Human Interaction Recognition Using Cross-modal Integration of Large Language Models
Understanding human-to-human interactions, especially in contexts like public security surveillance, is critical for monitoring and maintaining safety. Traditional activity recognition systems are limited by fixed vocabu…
Activity RecognitionHuman Interaction RecognitionVideo UnderstandingCHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action Recognition
Skeleton-based multi-entity action recognition is a challenging task aiming to identify interactive actions or group activities involving multiple diverse entities. Existing models for individuals often fall short in thi…
3D Action RecognitionAction RecognitionGroup Activity RecognitionHuman Interaction Recognition+1Empathic Grounding: Explorations using Multimodal Interaction and Large Language Models with Conversational Agents
We introduce the concept of "empathic grounding" in conversational agents as an extension of Clark's conceptualization of grounding in conversation in which the grounding criterion includes listener empathy for the speak…
Emotional IntelligenceEmotion ClassificationHuman Interaction RecognitionLanguage Modelling+4Exploring Vision Transformers for 3D Human Motion-Language Models with Motion Patches
To build a cross-modal latent space between 3D human motion and language, acquiring large-scale and high-quality human motion data is crucial. However, unlike the abundance of image data, the scarcity of motion data has …
Human Interaction RecognitionTransfer Learning