paper-with-me

Papers

Learning to Generate Long-term Future Narrations Describing Activities of Daily Living

2025-03-03 · Ramanathan Rajendiran, Debaditya Roy, Basura Fernando

Anticipating future events is crucial for various application domains such as healthcare, smart home technology, and surveillance. Narrative event descriptions provide context-rich information, enhancing a system's future planning and decision-making capabilities. We propose a novel task: $\textit{long-term future narration generation}$, which extends beyond traditional action anticipation by generating detailed narrations of future daily activities. We introduce a visual-language model, ViNa, specifically designed to address this challenging task. ViNa integrates long-term videos and corresponding narrations to generate a sequence of future narrations that predict subsequent events and actions over extended time horizons. ViNa extends existing multimodal models that perform only short-term predictions or describe observed videos by generating long-term future narrations for a broader range of daily activities. We also present a novel downstream application that leverages the generated narrations called future video retrieval to help users improve planning for a task by visualizing the future. We evaluate future narration generation on the largest egocentric dataset Ego4D.

📄 PDF Abstract BibTeX arXiv:2503.01416

Code (0)

등록된 구현이 없습니다.

Tasks

Action AnticipationDecision MakingLanguage ModelingLanguage ModellingVideo Retrieval

Similar Papers 제목 키워드 기반

Learning Video Representations from Large Language Models

2022-12-08 · CVPR 2023 1 · Yue Zhao, Ishan Misra, Philipp Krähenbühl, Rohit Girdhar

We introduce LaViLa, a new approach to learning video-language representations by leveraging Large Language Models (LLMs). We repurpose pre-trained LLMs to be conditioned on visual input, and finetune them to create auto…

Action ClassificationAction RecognitionDiversityEgocentric Activity Recognition+2

Learning to Localize Actions in Instructional Videos with LLM-Based Multi-Pathway Text-Video Alignment

2024-09-22 · Yuxiao Chen, Kai Li, Wentao Bao, Deep Patel 외

Learning to localize temporal boundaries of procedure steps in instructional videos is challenging due to the limited availability of annotated large-scale training videos. Recent works focus on learning the cross-modal …

Contrastive Learningcross-modal alignmentSemantic SimilaritySemantic Textual Similarity+2

TeaserGen: Generating Teasers for Long Documentaries

2024-10-08 · Weihan Xu, Paul Pu Liang, Haven Kim, Julian McAuley 외

Teasers are an effective tool for promoting content in entertainment, commercial and educational fields. However, creating an effective teaser for long videos is challenging for it requires long-range multimodal modeling…

Language ModellingLarge Language Model

Verified Training for Counterfactual Explanation Robustness under Data Shift

2024-03-06 · Anna P. Meyer, Yuhao Zhang, Aws Albarghouthi, Loris D'Antoni

Counterfactual explanations (CEs) enhance the interpretability of machine learning models by describing what changes to an input are necessary to change its prediction to a desired class. These explanations are commonly …

counterfactualCounterfactual Explanation

What You Say Is What You Show: Visual Narration Detection in Instructional Videos

2023-01-05 · Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, Kristen Grauman

Narrated ''how-to'' videos have emerged as a promising data source for a wide range of learning problems, from learning visual representations to training robot policies. However, this data is extremely noisy, as the nar…