paper-with-me

홈 › Papers

Multimodal Intent Discovery from Livestream Videos

2022-07-01 · Findings (NAACL) 2022 7 · Adyasha Maharana, Quan Tran, Franck Dernoncourt, Seunghyun Yoon, Trung Bui, Walter Chang, Mohit Bansal

Individuals, educational institutions, and businesses are prolific at generating instructional video content such as “how-to” and tutorial guides. While significant progress has been made in basic video understanding tasks, identifying procedural intent within these instructional videos is a challenging and important task that remains unexplored but essential to video summarization, search, and recommendations. This paper introduces the problem of instructional intent identification and extraction from software instructional livestreams. We construct and present a new multimodal dataset consisting of software instructional livestreams and containing manual annotations for both detailed and abstract procedural intent that enable training and evaluation of joint video and text understanding models. We then introduce a multimodal cascaded cross-attention model to efficiently combine the weaker and noisier video signal with the more discriminative text signal. Our experiments show that our proposed model brings significant gains compared to strong baselines, including large-scale pretrained multimodal models. Our analysis further identifies that the task benefits from spatial as well as motion features extracted from videos, and provides insight on how the video signal is preferentially used for intent discovery. We also show that current models struggle to comprehend the nature of abstract intents, revealing important gaps in multimodal understanding and paving the way for future work.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Intent DiscoveryVideo SummarizationVideo Understanding

Similar Papers 제목 키워드 기반

LiveSeg: Unsupervised Multimodal Temporal Segmentation of Long Livestream Videos

2022-10-12 · JieLin Qiu, Franck Dernoncourt, Trung Bui, Zhaowen Wang 외

Livestream videos have become a significant part of online learning, where design, digital marketing, creative painting, and other skills are taught by experienced experts in the sessions, making them valuable materials.…

MarketingSegmentation

FLUID: From Ephemeral IDs to Multimodal Semantic Codes for Industrial-Scale Livestreaming Recommendation

2026-05-20 · Xinhang Yuan, Zexi Huang, Anjia Cao, Xudong Lu 외 arxiv

Modern recommender systems rely heavily on ID-based collaborative filtering: each item is represented by a unique ID embedding that accumulates collaborative signals from user interactions. Livestreaming recommendation, …

Collaborative Filtering

LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding

2026-01-21 · Xiaodong Wang, Langling Huang, Zhirong Wu, Xu Zhao 외 arxiv

The development of multimodal large language models (MLLMs) has advanced general video understanding. However, existing video evaluation benchmarks primarily focus on non-interactive videos, such as movies and recordings…

BehanceCC: A ChitChat Detection Dataset For Livestreaming Video Transcripts

2022-06-01 · LREC 2022 6 · Viet Lai, Amir Pouran Ben Veyseh, Franck Dernoncourt, Thien Nguyen

Livestreaming videos have become an effective broadcasting method for both video sharing and educational purposes. However, livestreaming videos contain a considerable amount of off-topic content (i.e., up to 50%) which …

BehancePR: A Punctuation Restoration Dataset for Livestreaming Video Transcript

2022-07-01 · Findings (NAACL) 2022 7 · Viet Lai, Amir Pouran Ben Veyseh, Franck Dernoncourt, Thien Nguyen

Given the increasing number of livestreaming videos, automatic speech recognition and post-processing for livestreaming video transcripts are crucial for efficient data management as well as knowledge mining. A key step …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ManagementPunctuation Restoration+3