Action Understanding
1개 벤치마크 · 논문 141편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Prompted Contrast with Masked Motion Modeling: Towards Versatile 3D Action Representation Learning
LLaVA-Pose: Enhancing Human Pose and Action Understanding via Keypoint-Integrated Instruction Tuning
LLaVAction: evaluating and training multi-modal large language models for action recognition
Papers
MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching
Existing action quality assessment (AQA) datasets and methods rely primarily on visual inputs such as RGB and pose, overlooking physiological dynamics such as muscle mechanics and often modeling actions as monolithic pat…
Action Quality AssessmentAction UnderstandingRecognition-Conditioned Reasoning: A Training-Free Multimodal-LLM Pipeline for Fine-Grained Micro-Action Understanding
Micro-actions are subtle, short, low-amplitude body movements, such as a fidgeting hand or a slight head tilt, that humans perform with little conscious intent yet that reliably leak emotional and psychological state. Un…
Action UnderstandingG3Ego: Gaze-Guided Graphs for Egocentric Action Understanding
Egocentric action understanding is often addressed using large video models pretrained on extensive exocentric datasets. However, many first-person actions depend on a small number of hand-object interactions involving o…
Action UnderstandingAction RecognitionCompositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos
Assembly action understanding is a key enabler for effective human-robot collaborative assembly, yet it remains challenging due to subtle motions and fine-grained hand-object interactions. We adapt vision-language models…
Hyperparameter OptimizationAction UnderstandingMulti-Task LearningAction RecognitionPartial Skeleton Visibility for Action Recognition: A Constrained Field-of-View Approach
Skeleton-based action recognition has achieved remarkable success by exploiting joint coordinates and their topological connections, yet prevailing methods overwhelmingly assume complete and clean skeleton inputs. In rea…
Action UnderstandingAction RecognitionGold Points Sniper: Self-guided Visual Reasoning in VLM for Fine-grained Action Understanding
Robots operating in everyday environments must understand fine-grained human actions, intentions, and contextual cues from broad views where people occupy only small regions, a capability unmet by current systems. While …
Action UnderstandingMultimodal ReasoningAction RecognitionVisual Reasoning