paper-with-me

홈 › Papers

TV-Dialogue: Crafting Theme-Aware Video Dialogues with Immersive Interaction

2025-01-31 · Sai Wang, Fan Ma, Xinyi Li, Hehe Fan, Yu Wu

Recent advancements in LLMs have accelerated the development of dialogue generation across text and images, yet video-based dialogue generation remains underexplored and presents unique challenges. In this paper, we introduce Theme-aware Video Dialogue Crafting (TVDC), a novel task aimed at generating new dialogues that align with video content and adhere to user-specified themes. We propose TV-Dialogue, a novel multi-modal agent framework that ensures both theme alignment (i.e., the dialogue revolves around the theme) and visual consistency (i.e., the dialogue matches the emotions and behaviors of characters in the video) by enabling real-time immersive interactions among video characters, thereby accurately understanding the video content and generating new dialogue that aligns with the given themes. To assess the generated dialogues, we present a multi-granularity evaluation benchmark with high accuracy, interpretability and reliability, demonstrating the effectiveness of TV-Dialogue on self-collected dataset over directly using existing LLMs. Extensive experiments reveal that TV-Dialogue can generate dialogues for videos of any length and any theme in a zero-shot manner without training. Our findings underscore the potential of TV-Dialogue for various applications, such as video re-creation, film dubbing and its use in downstream multimodal tasks.

📄 PDF Abstract BibTeX arXiv:2501.18940

Code (0)

등록된 구현이 없습니다.

Tasks

Dialogue Generation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

CATCH: A Controllable Theme Detection Framework with Contextualized Clustering and Hierarchical Generation

2025-12-25 · Rui Ke, Jiahui Xu, Shenghao Yang, Kuang Wang 외 arxiv

Theme detection is a fundamental task in user-centric dialogue systems, aiming to identify the latent topic of each utterance without relying on predefined schemas. Unlike intent induction, which operates within fixed la…

$C^3$: Compositional Counterfactual Contrastive Learning for Video-grounded Dialogues

2021-06-16 · Hung Le, Nancy F. Chen, Steven C. H. Hoi

Video-grounded dialogue systems aim to integrate video understanding and dialogue understanding to generate responses that are relevant to both the dialogue and video context. Most existing approaches employ deep learnin…

Contrastive LearningcounterfactualDialogue UnderstandingMultimodal Reasoning+1

VSTAR: A Video-grounded Dialogue Dataset for Situated Semantic Understanding with Scene and Topic Transitions

2023-05-30 · Yuxuan Wang, Zilong Zheng, Xueliang Zhao, Jinpeng Li 외

Video-grounded dialogue understanding is a challenging problem that requires machine to perceive, parse and reason over situated semantics extracted from weakly aligned video and dialogues. Most existing benchmarks treat…

Dialogue GenerationDialogue UnderstandingScene SegmentationSegmentation

VScript: Controllable Script Generation with Visual Presentation

2022-03-01 · Ziwei Ji, Yan Xu, I-Tsun Cheng, Samuel Cahyawijaya 외

In order to offer a customized script tool and inspire professional scriptwriters, we present VScript. It is a controllable pipeline that generates complete scripts, including dialogues and scene descriptions, as well as…

Dialogue GenerationRetrievalScript GenerationVideo Retrieval

Video-Grounded Dialogues with Pretrained Generation Language Models

2020-06-27 · ACL 2020 6 · Hung Le, Steven C. H. Hoi

Pre-trained language models have shown remarkable success in improving various downstream NLP tasks due to their ability to capture dependencies in textual data and generate natural responses. In this paper, we leverage …

Sentence