Learning to Generate Long-term Future Narrations Describing Activities of Daily Living
Anticipating future events is crucial for various application domains such as healthcare, smart home technology, and surveillance. Narrative event descriptions provide context-rich information, enhancing a system's future planning and decision-making capabilities. We propose a novel task: $\textit{long-term future narration generation}$, which extends beyond traditional action anticipation by generating detailed narrations of future daily activities. We introduce a visual-language model, ViNa, specifically designed to address this challenging task. ViNa integrates long-term videos and corresponding narrations to generate a sequence of future narrations that predict subsequent events and actions over extended time horizons. ViNa extends existing multimodal models that perform only short-term predictions or describe observed videos by generating long-term future narrations for a broader range of daily activities. We also present a novel downstream application that leverages the generated narrations called future video retrieval to help users improve planning for a task by visualizing the future. We evaluate future narration generation on the largest egocentric dataset Ego4D.
Code (0)
등록된 구현이 없습니다.
Tasks
Action AnticipationDecision MakingLanguage ModelingLanguage ModellingVideo RetrievalSimilar Papers 제목 키워드 기반
Learning Video Representations from Large Language Models
We introduce LaViLa, a new approach to learning video-language representations by leveraging Large Language Models (LLMs). We repurpose pre-trained LLMs to be conditioned on visual input, and finetune them to create auto…
Action ClassificationAction RecognitionDiversityEgocentric Activity Recognition+2Learning to Localize Actions in Instructional Videos with LLM-Based Multi-Pathway Text-Video Alignment
Learning to localize temporal boundaries of procedure steps in instructional videos is challenging due to the limited availability of annotated large-scale training videos. Recent works focus on learning the cross-modal …
Contrastive Learningcross-modal alignmentSemantic SimilaritySemantic Textual Similarity+2TeaserGen: Generating Teasers for Long Documentaries
Teasers are an effective tool for promoting content in entertainment, commercial and educational fields. However, creating an effective teaser for long videos is challenging for it requires long-range multimodal modeling…
Language ModellingLarge Language ModelVerified Training for Counterfactual Explanation Robustness under Data Shift
Counterfactual explanations (CEs) enhance the interpretability of machine learning models by describing what changes to an input are necessary to change its prediction to a desired class. These explanations are commonly …
counterfactualCounterfactual ExplanationWhat You Say Is What You Show: Visual Narration Detection in Instructional Videos
Narrated ''how-to'' videos have emerged as a promising data source for a wide range of learning problems, from learning visual representations to training robot policies. However, this data is extremely noisy, as the nar…