paper-with-me

Action Segmentation

9개 벤치마크 · 논문 263편 · 이 태스크의 논문 보기 →

Benchmarks

Breakfast

결과 74개

50 Salads

결과 56개

GTEA

결과 56개

COIN

결과 18개

Assembly101

결과 14개

JIGSAWS

결과 14개

50Salads

결과 2개

Most implemented

Papers

Motion-Saliency Complementary Masked Modeling for Point Cloud Video Understanding

2026-08-31 · Wei Wang, Yiding Sun, Yuyan Wang, Zhuoyue Zhang 외 arxiv

Point cloud video representation learning is crucial for 3D dynamic scene understanding. In this paper, we propose MoSaiC, a novel Motion-Saliency Complementary masked modeling framework for self-supervised point cloud v…

Representation LearningSemantic SegmentationScene UnderstandingAction Segmentation

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

2026-07-15 · Zishuo Li, Bowen Yang, Changtao Miao, Kai Zhu 외 hf

Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable t…

Action Segmentation

Adaptive Latent Trajectory Anchoring for Action Segmentation Dataset Condensation

2026-07-10 · Artheme Gauthier-Villar, Guodong Ding, Angela Yao arxiv

Dataset condensation for action segmentation synthesizes compact, informative representations of long, untrimmed video datasets. The existing approach relies on Variational Autoencoders and an iterative latent optimizati…

Action Segmentation

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture

2026-06-22 · Felix Tristram, Stefano Gasperini, Benjamin Killeen, Marcel Walch 외 arxiv

The increasing maturity of embodied AI platforms has driven a growing interest in procedural video representation learning to support intelligent assistance systems for complex, multi-step tasks. Leveraging large-scale l…

Representation LearningAction ClassificationAction Segmentation

The Unreasonable Effectiveness of VLMs for Zero-shot Procedural Mistake Detection

2026-06-19 · Serdar Ozsoy, Lars Doorenbos, Federico Spurio, Gianpiero Francesca 외 arxiv

Procedural mistake detection is important for quality control and user assistance across many disciplines. Recent work in this field has achieved significant gains by using the reasoning capabilities of Video-Language Mo…

Action Segmentation

UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning

2026-06-18 · Wenhao Chi, Arkaprava Sinha, Dominick Reilly, Hieu Le 외 arxiv

Egocentric video understanding is inherently limited by the narrow perspective of wearable cameras: a single viewpoint, a single modality, a single model cannot capture the full richness of human action. We argue that a …

Representation LearningAction SegmentationAction RecognitionVideo Retrieval

전체 263편 보기 →