paper-with-me

Papers

Grounding Foundational Vision Models with 3D Human Poses for Robust Action Recognition

2025-11-06 · Nicholas Babey, Tiffany Gu, Yiheng Li, Cristian Meo, Kevin Zhu arxiv

For embodied agents to effectively understand and interact within the world around them, they require a nuanced comprehension of human actions grounded in physical space. Current action recognition models, often relying on RGB video, learn superficial correlations between patterns and action labels, so they struggle to capture underlying physical interaction dynamics and human poses in complex scenes. We propose a model architecture that grounds action recognition in physical space by fusing two powerful, complementary representations: V-JEPA 2's contextual, predictive world dynamics and CoMotion's explicit, occlusion-tolerant human pose data. Our model is validated on both the InHARD and UCF-19-Y-OCC benchmarks for general action recognition and high-occlusion action recognition, respectively. Our model outperforms three other baselines, especially within complex, occlusive scenes. Our findings emphasize a need for action recognition to be supported by spatial understanding instead of statistical pattern recognition.

📄 PDF Abstract BibTeX arXiv:2511.05622

Code (0)

등록된 구현이 없습니다.

Tasks

Action Recognition

Similar Papers 제목 키워드 기반

Large language models for artificial general intelligence (AGI): A survey of foundational principles and approaches

2025-01-06 · Alhassan Mumuni, Fuseini Mumuni

Generative artificial intelligence (AI) systems based on large-scale pretrained foundation models (PFMs) such as vision-language models, large language models (LLMs), diffusion models and vision-language-action (VLA) mod…

Vision-Language-Action

Interacted Object Grounding in Spatio-Temporal Human-Object Interactions

2024-12-27 · Xiaoyang Liu, Boran Wen, Xinpeng Liu, Zizheng Zhou 외

Spatio-temporal Human-Object Interaction (ST-HOI) understanding aims at detecting HOIs from videos, which is crucial for activity understanding. However, existing whole-body-object interaction video benchmarks overlook t…

Human-Object Interaction DetectionObjectQuestion Answering

Meanings and Measurements: Multi-Agent Probabilistic Grounding for Vision-Language Navigation

2026-03-19 · Swagat Padhan, Lakshya Jain, Bhavya Minesh Shah, Omkar Patil 외 arxiv

Robots collaborating with humans must convert natural language goals into actionable, physically grounded decisions. For example, executing a command such as "go two meters to the right of the fridge" requires grounding …

Vision-Language Navigation

Green-VLA: Staged Vision-Language-Action Model for Generalist Robots

2026-01-31 · I. Apanasevich, M. Artemyev, R. Babakyan, P. Fedotova 외 arxiv

We introduce Green-VLA, a staged Vision-Language-Action (VLA) framework for real-world deployment on the Green humanoid robot while maintaining generalization across diverse embodiments. Green-VLA follows a five stage cu…

Out-of-Distribution Detection

MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning

2025-07-19 · Liujian Tang, Shaokang Dong, Yijia Huang, Minqi Xiang 외 arxiv

This paper presents MagicGUI, a foundational mobile GUI agent designed to address critical challenges in perception, grounding, and reasoning within real-world mobile GUI environments. The framework is underpinned by fol…