paper-with-me

홈 › Papers

A Closer Look at Temporal Ordering in the Segmentation of Instructional Videos

2022-09-30 · Anil Batra, Shreyank N Gowda, Frank Keller, Laura Sevilla-Lara

Understanding the steps required to perform a task is an important skill for AI systems. Learning these steps from instructional videos involves two subproblems: (i) identifying the temporal boundary of sequentially occurring segments and (ii) summarizing these steps in natural language. We refer to this task as Procedure Segmentation and Summarization (PSS). In this paper, we take a closer look at PSS and propose three fundamental improvements over current methods. The segmentation task is critical, as generating a correct summary requires each step of the procedure to be correctly identified. However, current segmentation metrics often overestimate the segmentation quality because they do not consider the temporal order of segments. In our first contribution, we propose a new segmentation metric that takes into account the order of segments, giving a more reliable measure of the accuracy of a given predicted segmentation. Current PSS methods are typically trained by proposing segments, matching them with the ground truth and computing a loss. However, much like segmentation metrics, existing matching algorithms do not consider the temporal order of the mapping between candidate segments and the ground truth. In our second contribution, we propose a matching algorithm that constrains the temporal order of segment mapping, and is also differentiable. Lastly, we introduce multi-modal feature training for PSS, which further improves segmentation. We evaluate our approach on two instructional video datasets (YouCook2 and Tasty) and observe an improvement over the state-of-the-art of $\sim7\%$ and $\sim2.5\%$ for procedure segmentation and summarization, respectively.

📄 PDF Abstract BibTeX arXiv:2209.15501

Code (0)

등록된 구현이 없습니다.

Tasks

Dense Video CaptioningSegmentation

Similar Papers 제목 키워드 기반

Learning Procedure-aware Video Representation from Instructional Videos and Their Narrations

2023-03-31 · CVPR 2023 1 · Yiwu Zhong, Licheng Yu, Yang Bai, Shangwen Li 외

The abundance of instructional videos and their narrations over the Internet offers an exciting avenue for understanding procedural activities. In this work, we propose to learn video representation that encodes both act…

Investigating the Reordering Capability in CTC-based Non-Autoregressive End-to-End Speech Translation

2021-05-11 · Findings (ACL) 2021 8 · Shun-Po Chuang, Yung-Sung Chuang, Chih-Chiang Chang, Hung-Yi Lee

We study the possibilities of building a non-autoregressive speech-to-text translation model using connectionist temporal classification (CTC), and use CTC-based automatic speech recognition as an auxiliary task to impro…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+3

Long-Tail Temporal Action Segmentation with Group-wise Temporal Logit Adjustment

2024-08-19 · Zhanzhong Pang, Fadime Sener, Shrinivas Ramasubramanian, Angela Yao

Procedural activity videos often exhibit a long-tailed action distribution due to varying action frequencies and durations. However, state-of-the-art temporal action segmentation methods overlook the long tail and fail t…

Action SegmentationSegmentationTemporal Action Segmentation

Unsupervised Action Segmentation for Instructional Videos

2021-06-07 · AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo, Irfan Essa

In this paper we address the problem of automatically discovering atomic actions in unsupervised manner from instructional videos, which are rarely annotated with atomic actions. We present an unsupervised approach to le…

Action SegmentationSegmentationUnsupervised Action Segmentation

Comprehensive Instructional Video Analysis: The COIN Dataset and Performance Evaluation

2020-03-20 · Yansong Tang, Jiwen Lu, Jie zhou

Thanks to the substantial and explosively inscreased instructional videos on the Internet, novices are able to acquire knowledge for completing various tasks. Over the past decade, growing efforts have been devoted to in…

Action Detection