Detours for Navigating Instructional Videos
We introduce the video detours problem for navigating instructional videos. Given a source video and a natural language query asking to alter the how-to video's current path of execution in a certain way, the goal is to find a related ''detour video'' that satisfies the requested alteration. To address this challenge, we propose VidDetours, a novel video-language approach that learns to retrieve the targeted temporal segments from a large repository of how-to's using video-and-text conditioned queries. Furthermore, we devise a language-based pipeline that exploits how-to video narration text to create weakly supervised training data. We demonstrate our idea applied to the domain of how-to cooking videos, where a user can detour from their current recipe to find steps with alternate ingredients, tools, and techniques. Validating on a ground truth annotated dataset of 16K samples, we show our model's significant improvements over best available methods for video retrieval and question answering, with recall rates exceeding the state of the art by 35%.
Code (0)
등록된 구현이 없습니다.
Tasks
16kQuestion AnsweringRetrievalVideo RetrievalSimilar Papers 제목 키워드 기반
Unsupervised Discovery of Actions in Instructional Videos
In this paper we address the problem of automatically discovering atomic actions in unsupervised manner from instructional videos. Instructional videos contain complex activities and are a rich source of information for …
Collaborative Weakly Supervised Video Correlation Learning for Procedure-Aware Instructional Video Analysis
Video Correlation Learning (VCL), which aims to analyze the relationships between videos, has been widely studied and applied in various general video tasks. However, applying VCL to instructional videos is still quite c…
Action Quality AssessmentProcedure LearningBeyond Instructional Videos: Probing for More Diverse Visual-Textual Grounding on YouTube
Pretraining from unlabelled web videos has quickly become the de-facto means of achieving high performance on many video understanding tasks. Features are learned via prediction of grounded relationships between visual c…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1Unsupervised Action Segmentation for Instructional Videos
In this paper we address the problem of automatically discovering atomic actions in unsupervised manner from instructional videos, which are rarely annotated with atomic actions. We present an unsupervised approach to le…
Action SegmentationSegmentationUnsupervised Action SegmentationHow to Make a BLT Sandwich? Learning to Reason towards Understanding Web Instructional Videos
Understanding web instructional videos is an essential branch of video understanding in two aspects. First, most existing video methods focus on short-term actions for a-few-second-long video clips; these methods are not…
Logical ReasoningQuestion AnsweringVideo Understanding