paper-with-me

Papers

NarraScore: Bridging Visual Narrative and Musical Dynamics via Hierarchical Affective Control

2026-02-09 · Yufan Wen, Zhaocheng Liu, YeGuo Hua, Ziyi Guo, Lihua Zhang, Chun Yuan, Jian Wu arxiv

Synthesizing coherent soundtracks for long-form videos remains a formidable challenge, currently stalled by three critical impediments: computational scalability, temporal coherence, and, most critically, a pervasive semantic blindness to evolving narrative logic. To bridge these gaps, we propose NarraScore, a hierarchical framework predicated on the core insight that emotion serves as a high-density compression of narrative logic. Uniquely, we repurpose frozen Vision-Language Models (VLMs) as continuous affective sensors, distilling high-dimensional visual streams into dense, narrative-aware Valence-Arousal trajectories. Mechanistically, NarraScore employs a Dual-Branch Injection strategy to reconcile global structure with local dynamism: a \textit{Global Semantic Anchor} ensures stylistic stability, while a surgical \textit{Token-Level Affective Adapter} modulates local tension via direct element-wise residual injection. This minimalist design bypasses the bottlenecks of dense attention and architectural cloning, effectively mitigating the overfitting risks associated with data scarcity. Experiments demonstrate that NarraScore achieves state-of-the-art consistency and narrative alignment with negligible computational overhead, establishing a fully autonomous paradigm for long-video soundtrack generation.

📄 PDF Abstract BibTeX arXiv:2602.09070

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Interactive Narrative Analytics: Bridging Computational Narrative Extraction and Human Sensemaking

2026-01-16 · Brian Keith arxiv

Information overload and misinformation create significant challenges in extracting meaningful narratives from large news collections. This paper defines the nascent field of Interactive Narrative Analytics (INA), which …

GANterpretations

2020-11-06 · Pablo Samuel Castro

Since the introduction of Generative Adversarial Networks (GANs) [Goodfellow et al., 2014] there has been a regular stream of both technical advances (e.g., Arjovsky et al. [2017]) and creative uses of these generative m…

VisTopics: A Visual Semantic Unsupervised Approach to Topic Modeling of Video and Image Data

2025-05-20 · Ayse D Lokmanoglu, Dror Walter

Understanding visual narratives is crucial for examining the evolving dynamics of media representation. This study introduces VisTopics, a computational framework designed to analyze large-scale visual datasets through a…

Generative Artificial Intelligence, Musical Heritage and the Construction of Peace Narratives: A Case Study in Mali

2026-01-21 · Nouhoum Coulibaly, Ousmane Ly, Michael Leventhal, Ousmane Goro arxiv

This study explores the capacity of generative artificial intelligence (Gen AI) to contribute to the construction of peace narratives and the revitalization of musical heritage in Mali. The study has been made in a polit…

Musical Word Embedding: Bridging the Gap between Listening Contexts and Music

2020-07-23 · Seungheon Doh, Jongpil Lee, Tae Hong Park, Juhan Nam

Word embedding pioneered by Mikolov et al. is a staple technique for word representations in natural language processing (NLP) research which has also found popularity in music information retrieval tasks. Depending on t…

Information RetrievalMusic Information RetrievalRetrieval