paper-with-me

Papers

Video-Mined Task Graphs for Keystep Recognition in Instructional Videos

2023-07-17 · NeurIPS 2023 11

Procedural activity understanding requires perceiving human actions in terms of a broader task, where multiple keysteps are performed in sequence across a long video to reach a final goal state -- such as the steps of a recipe or a DIY fix-it task. Prior work largely treats keystep recognition in isolation of this broader structure, or else rigidly confines keysteps to align with a predefined sequential script. We propose discovering a task graph automatically from how-to videos to represent probabilistically how people tend to execute keysteps, and then leverage this graph to regularize keystep recognition in novel videos. On multiple datasets of real-world instructional videos, we show the impact: more reliable zero-shot keystep localization and improved video representation learning, exceeding the state of the art.

📄 PDF Abstract BibTeX arXiv:2307.08763

Code (1)

facebookresearch/taskgraph pytorch

Tasks

Representation Learning

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Graph-Based Multimodal and Multi-view Alignment for Keystep Recognition

2025-01-07 · Julia Lee Romero, Kyle Min, Subarna Tripathi, Morteza Karimzadeh

Egocentric videos capture scenes from a wearer's viewpoint, resulting in dynamic backgrounds, frequent motion, and occlusions, posing challenges to accurate keystep recognition. We propose a flexible graph-learning frame…

Graph LearningNode Classification

MistExit: Learning to Exit for Early Mistake Detection in Procedural Videos

2026-03-15 · Sagnik Majumder, Anish Nethi, Ziad Al-Halah, Kristen Grauman arxiv

We introduce the task of early mistake detection in video, where the goal is to determine whether a keystep in a procedural activity is performed correctly while observing as little of the streaming video as possible. To…

Reinforcement Learning

ENIGMA-360: An Ego-Exo Dataset for Human Behavior Understanding in Industrial Scenarios

2026-03-10 · Francesco Ragusa, Rosario Leonardi, Michele Mazzamuto, Daniele Di Mauro 외 arxiv

Understanding human behavior from complementary egocentric (ego) and exocentric (exo) points of view enables the development of systems that can support workers in industrial environments and enhance their safety. Howeve…

Human-Object Interaction DetectionAction Segmentation

EgoEMS: A High-Fidelity Multimodal Egocentric Dataset for Cognitive Assistance in Emergency Medical Services

2025-11-13 · Keshara Weerasinghe, Xueren Ge, Tessa Heick, Lahiru Nuwan Wijayasingha 외 arxiv

Emergency Medical Services (EMS) are critical to patient survival in emergencies, but first responders often face intense cognitive demands in high-stakes situations. AI cognitive assistants, acting as virtual partners, …

Speaker DiarizationDecision Making

Personalized Cinemagraphs using Semantic Understanding and Collaborative Learning

2017-08-09 · ICCV 2017 10 · Tae-Hyun Oh, Kyungdon Joo, Neel Joshi, Baoyuan Wang 외

Cinemagraphs are a compelling way to convey dynamic aspects of a scene. In these media, dynamic and still elements are juxtaposed to create an artistic and narrative experience. Creating a high-quality, aesthetically ple…

Object RecognitionSemantic Segmentation