Action Parsing
1개 벤치마크 · 논문 18편 · 이 태스크의 논문 보기 →
Benchmarks
JerichoWorld
Most implemented
Modeling Worlds in Text
Frontal Low-rank Random Tensors for Fine-grained Action Segmentation
Local Temporal Bilinear Pooling for Fine-grained Action Parsing
Papers
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
Towards an embodied generalist for real-world interaction, Multimodal Large Language Model (MLLM) agents still suffer from challenging latency, sparse feedback, and irreversible mistakes. Video games offer an ideal testb…
Action ParsingInstruct2Act: From Human Instruction to Actions Sequencing and Execution via Robot Action Network for Robotic Manipulation
Robots often struggle to follow free-form human instructions in real-world settings due to computational and sensing limitations. We address this gap with a lightweight, fully on-device pipeline that converts natural-lan…
Action ParsingContext Matters: Peer-Aware Student Behavioral Engagement Measurement via VLM Action Parsing and LLM Sequence Classification
Understanding student behavior in the classroom is essential to improve both pedagogical quality and student engagement. Existing methods for predicting student engagement typically require substantial annotated data to …
Action RecognitionAction ParsingPIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic Manipulation
Language-guided robotic manipulation is a challenging task that requires an embodied agent to follow abstract user instructions to accomplish various complex manipulation tasks. Previous work trivially fitting the data w…
Action ParsingAction parsing using context features
We propose an action parsing algorithm to parse a video sequence containing an unknown number of actions into its action segments. We argue that context information, particularly the temporal information about other acti…
Action ParsingAction SegmentationSegmentationPart-level Action Parsing via a Pose-guided Coarse-to-Fine Framework
Action recognition from videos, i.e., classifying a video into one of the pre-defined action types, has been a popular topic in the communities of artificial intelligence, multimedia, and signal processing. However, exis…
Action ParsingAction Recognition