paper-with-me

Papers

Facilitating Video Story Interaction with Multi-Agent Collaborative System

2025-05-02 · Yiwen Zhang, Jianing Hao, Zhan Wang, Hongling Sheng, Wei Zeng

Video story interaction enables viewers to engage with and explore narrative content for personalized experiences. However, existing methods are limited to user selection, specially designed narratives, and lack customization. To address this, we propose an interactive system based on user intent. Our system uses a Vision Language Model (VLM) to enable machines to understand video stories, combining Retrieval-Augmented Generation (RAG) and a Multi-Agent System (MAS) to create evolving characters and scene experiences. It includes three stages: 1) Video story processing, utilizing VLM and prior knowledge to simulate human understanding of stories across three modalities. 2) Multi-space chat, creating growth-oriented characters through MAS interactions based on user queries and story stages. 3) Scene customization, expanding and visualizing various story scenes mentioned in dialogue. Applied to the Harry Potter series, our study shows the system effectively portrays emergent character social behavior and growth, enhancing the interactive experience in the video story world.

📄 PDF Abstract BibTeX arXiv:2505.03807

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingRAGRetrieval-augmented Generation

Methods 이 논문이 사용한 방법론

MAS This optimizer mix ADAM and SGD creating the MAS optimizer.

Similar Papers 제목 키워드 기반

StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration

2024-11-07 · Panwen Hu, Jin Jiang, Jianqi Chen, Mingfei Han 외

The advent of AI-Generated Content (AIGC) has spurred research into automated video generation to streamline conventional processes. However, automating storytelling video production, particularly for customized narrativ…

Video Generation

Modeling User Empathy Elicited by a Robot Storyteller

2021-07-29 · Leena Mathur, Micol Spitale, Hao Xi, Jieyun Li 외

Virtual and robotic agents capable of perceiving human empathy have the potential to participate in engaging and meaningful human-machine interactions that support human well-being. Prior research in computational empath…

FriendsQA: A New Large-Scale Deep Video Understanding Dataset with Fine-grained Topic Categorization for Story Videos

2024-12-22 · Zhengqian Wu, Ruizhe Li, Zijun Xu, Zhongyuan Wang 외

Video question answering (VideoQA) aims to answer natural language questions according to the given videos. Although existing models perform well in the factoid VideoQA task, they still face challenges in deep video unde…

Language ModellingLarge Language ModelQuestion AnsweringVideo Question Answering+1

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration

2026-04-06 · Zhongyu Yang, Zuhao Yang, Shuo Zhan, Tan Yue 외 arxiv

Video question answering (VideoQA) is a challenging task that requires integrating spatial, temporal, and semantic information to capture the complex dynamics of video sequences. Although recent advances have introduced …

Video Question Answering

Video Dialog via Progressive Inference and Cross-Transformer

2019-11-01 · IJCNLP 2019 11 · Weike Jin, Zhou Zhao, Mao Gu, Jun Xiao 외

Video dialog is a new and challenging task, which requires the agent to answer questions combining video information with dialog history. And different from single-turn video question answering, the additional dialog his…

Answer GenerationQuestion AnsweringQuestion GenerationQuestion-Generation+2