paper-with-me

Action Parsing

1개 벤치마크 · 논문 18편 · 이 태스크의 논문 보기 →

Benchmarks

JerichoWorld

결과 3개

Most implemented

Modeling Worlds in Text

2021-05-21 · 구현 1개

Papers

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents

2026-04-08 · Mingyu Ouyang, Siyuan Hu, Kevin Qinghong Lin, Hwee Tou Ng 외 arxiv

Towards an embodied generalist for real-world interaction, Multimodal Large Language Model (MLLM) agents still suffer from challenging latency, sparse feedback, and irreversible mistakes. Video games offer an ideal testb…

Action Parsing

Instruct2Act: From Human Instruction to Actions Sequencing and Execution via Robot Action Network for Robotic Manipulation

2026-02-10 · Archit Sharma, Dharmendra Sharma, John Rebeiro, Peeyush Thakur 외 arxiv

Robots often struggle to follow free-form human instructions in real-world settings due to computational and sensing limitations. We address this gap with a lightweight, fully on-device pipeline that converts natural-lan…

Action Parsing

Context Matters: Peer-Aware Student Behavioral Engagement Measurement via VLM Action Parsing and LLM Sequence Classification

2026-01-10 · Ahmed Abdelkawy, Ahmed Elsayed, Asem Ali, Aly Farag 외 arxiv

Understanding student behavior in the classroom is essential to improve both pedagogical quality and student engagement. Existing methods for predicting student engagement typically require substantial annotated data to …

Action RecognitionAction Parsing

PIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic Manipulation

2024-10-14 · Kaidong Zhang, Pengzhen Ren, Bingqian Lin, Junfan Lin 외

Language-guided robotic manipulation is a challenging task that requires an embodied agent to follow abstract user instructions to accomplish various complex manipulation tasks. Previous work trivially fitting the data w…

Action Parsing

Action parsing using context features

2022-05-20 · Nagita Mehrseresht

We propose an action parsing algorithm to parse a video sequence containing an unknown number of actions into its action segments. We argue that context information, particularly the temporal information about other acti…

Action ParsingAction SegmentationSegmentation

Part-level Action Parsing via a Pose-guided Coarse-to-Fine Framework

2022-03-09 · Xiaodong Chen, Xinchen Liu, Wu Liu, Kun Liu 외

Action recognition from videos, i.e., classifying a video into one of the pre-defined action types, has been a popular topic in the communities of artificial intelligence, multimedia, and signal processing. However, exis…

Action ParsingAction Recognition

전체 18편 보기 →