Learning a Grammar Inducer from Massive Uncurated Instructional Videos
Video-aided grammar induction aims to leverage video information for finding more accurate syntactic grammars for accompanying text. While previous work focuses on building systems for inducing grammars on text that are well-aligned with video content, we investigate the scenario, in which text and video are only in loose correspondence. Such data can be found in abundance online, and the weak correspondence is similar to the indeterminacy problem studied in language acquisition. Furthermore, we build a new model that can better learn video-span correlation without manually designed features adopted by previous work. Experiments show that our model trained only on large-scale YouTube data with no text-video alignment reports strong and robust performances across three unseen datasets, despite domain shift and noisy label issues. Furthermore our model yields higher F1 scores than the previous state-of-the-art systems trained on in-domain data.
Code (1)
Tasks
Language AcquisitionVideo AlignmentSimilar Papers 제목 키워드 기반
Depth-bounding is effective: Improvements and evaluation of unsupervised PCFG induction
There have been several recent attempts to improve the accuracy of grammar induction systems by bounding the recursive complexity of the induction model (Ponvert et al., 2011; Noji and Johnson, 2016; Shain et al., 2016; …
Self-Supervised Multi-Task Procedure Learning from Instructional Videos
We address the problem of unsupervised procedure learning from instructional videos of multiple tasks using Deep Neural Networks (DNNs). Unlike existing works, we assume that training videos come from multiple tasks with…
Procedure LearningVideo ClassificationEnd-to-End Learning of Visual Representations from Uncurated Instructional Videos
Annotating videos is cumbersome, expensive and not scalable. Yet, many strong video models still rely on manually annotated data. With the recent introduction of the HowTo100M dataset, narrated videos now offer the possi…
Action LocalizationAction RecognitionAction SegmentationLong Video Retrieval (Background Removed)+4The Importance of Category Labels in Grammar Induction with Child-directed Utterances
Recent progress in grammar induction has shown that grammar induction is possible without explicit assumptions of language-specific knowledge. However, evaluation of induced grammars usually has ignored phrasal labels, a…
TL;DW? Summarizing Instructional Videos with Task Relevance & Cross-Modal Saliency
YouTube users looking for instructions for a specific task may spend a long time browsing content trying to find the right video that matches their needs. Creating a visual summary (abridged version of a video) provides …
ArticlesVideo Summarization