Action Segmentation
9개 벤치마크 · 논문 263편 · 이 태스크의 논문 보기 →
Benchmarks
Breakfast
50 Salads
GTEA
COIN
Assembly101
JIGSAWS
50Salads
MPII Cooking 2 Dataset
Most implemented
Temporal Convolutional Networks for Action Segmentation and Detection
End-to-End Learning of Visual Representations from Uncurated Instructional Videos
Hierarchical NeuroSymbolic Approach for Comprehensive and Explainable Action Quality Assessment
Temporal Action Segmentation: An Analysis of Modern Techniques
LOGO: A Long-Form Video Dataset for Group Action Quality Assessment
Papers
Motion-Saliency Complementary Masked Modeling for Point Cloud Video Understanding
Point cloud video representation learning is crucial for 3D dynamic scene understanding. In this paper, we propose MoSaiC, a novel Motion-Saliency Complementary masked modeling framework for self-supervised point cloud v…
Representation LearningSemantic SegmentationScene UnderstandingAction SegmentationOpen-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable t…
Action SegmentationAdaptive Latent Trajectory Anchoring for Action Segmentation Dataset Condensation
Dataset condensation for action segmentation synthesizes compact, informative representations of long, untrimmed video datasets. The existing approach relies on Variational Autoencoders and an iterative latent optimizati…
Action SegmentationP-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture
The increasing maturity of embodied AI platforms has driven a growing interest in procedural video representation learning to support intelligent assistance systems for complex, multi-step tasks. Leveraging large-scale l…
Representation LearningAction ClassificationAction SegmentationThe Unreasonable Effectiveness of VLMs for Zero-shot Procedural Mistake Detection
Procedural mistake detection is important for quality control and user assistance across many disciplines. Recent work in this field has achieved significant gains by using the reasoning capabilities of Video-Language Mo…
Action SegmentationUNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning
Egocentric video understanding is inherently limited by the narrow perspective of wearable cameras: a single viewpoint, a single modality, a single model cannot capture the full richness of human action. We argue that a …
Representation LearningAction SegmentationAction RecognitionVideo Retrieval