paper-with-me

홈 › Papers

Generating Event-oriented Attribution for Movies via Two-Stage Prefix-Enhanced Multimodal LLM

2024-09-14 · Yuanjie Lyu, Tong Xu, Zihan Niu, Bo Peng, Jing Ke, Enhong Chen

The prosperity of social media platforms has raised the urgent demand for semantic-rich services, e.g., event and storyline attribution. However, most existing research focuses on clip-level event understanding, primarily through basic captioning tasks, without analyzing the causes of events across an entire movie. This is a significant challenge, as even advanced multimodal large language models (MLLMs) struggle with extensive multimodal information due to limited context length. To address this issue, we propose a Two-Stage Prefix-Enhanced MLLM (TSPE) approach for event attribution, i.e., connecting associated events with their causal semantics, in movie videos. In the local stage, we introduce an interaction-aware prefix that guides the model to focus on the relevant multimodal information within a single clip, briefly summarizing the single event. Correspondingly, in the global stage, we strengthen the connections between associated events using an inferential knowledge graph, and design an event-aware prefix that directs the model to focus on associated events rather than all preceding clips, resulting in accurate event attribution. Comprehensive evaluations of two real-world datasets demonstrate that our framework outperforms state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2409.09362

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

A dataset for Audio-Visual Sound Event Detection in Movies

2023-02-14 · Rajat Hebbar, Digbalay Bose, Krishna Somandepalli, Veena Vijai 외

Audio event detection is a widely studied audio processing task, with applications ranging from self-driving cars to healthcare. In-the-wild datasets such as Audioset have propelled research in this field. However, many …

Event DetectionSelf-Driving CarsSound ClassificationSound Event Detection

CoTAR: Chain-of-Thought Attribution Reasoning with Multi-level Granularity

2024-04-16 · Moshe Berchansky, Daniel Fleischer, Moshe Wasserblat, Peter Izsak

State-of-the-art performance in QA tasks is currently achieved by systems employing Large Language Models (LLMs), however these models tend to hallucinate information in their responses. One approach focuses on enhancing…

Question Answering

CHATTER: A Character Attribution Dataset for Narrative Understanding

2024-11-07 · Sabyasachee Baruah, Shrikanth Narayanan

Computational narrative understanding studies the identification, description, and interaction of the elements of a narrative: characters, attributes, events, and relations. Narrative research has given considerable atte…

Attribute

Grasp-Then-Plan with Failure Attribution: A Closed Two-Stage Framework for Precise and Generalizable Robotic Manipulation

2026-06-02 · Jiahao Xu, Peiyuan Wang, Hanzhuo Zhang, Zihao Yu 외 arxiv

In robotic manipulation, the tight coupling between grasping and motion planning often obscures the true source of failure, leading to inefficient trial-and-error. To enable efficient long-horizon manipulation, we propos…

Motion Planning

Hollywood Identity Bias Dataset: A Context Oriented Bias Analysis of Movie Dialogues

2022-05-31 · LREC 2022 6 · Sandhya Singh, Prapti Roy, Nihar Sahoo, Niteesh Mallela 외

Movies reflect society and also hold power to transform opinions. Social biases and stereotypes present in movies can cause extensive damage due to their reach. These biases are not always found to be the need of storyli…