paper-with-me

홈 › Papers

Learning to Highlight Audio by Watching Movies

2025-05-17 · CVPR 2025 1 · Chao Huang, Ruohan Gao, J. M. F. Tsang, Jan Kurcius, Cagdas Bilen, Chenliang Xu, Anurag Kumar, Sanjeel Parekh

Recent years have seen a significant increase in video content creation and consumption. Crafting engaging content requires the careful curation of both visual and audio elements. While visual cue curation, through techniques like optimal viewpoint selection or post-editing, has been central to media production, its natural counterpart, audio, has not undergone equivalent advancements. This often results in a disconnect between visual and acoustic saliency. To bridge this gap, we introduce a novel task: visually-guided acoustic highlighting, which aims to transform audio to deliver appropriate highlighting effects guided by the accompanying video, ultimately creating a more harmonious audio-visual experience. We propose a flexible, transformer-based multimodal framework to solve this task. To train our model, we also introduce a new dataset -- the muddy mix dataset, leveraging the meticulous audio and video crafting found in movies, which provides a form of free supervision. We develop a pseudo-data generation process to simulate poorly mixed audio, mimicking real-world scenarios through a three-step process -- separation, adjustment, and remixing. Our approach consistently outperforms several baselines in both quantitative and subjective evaluation. We also systematically study the impact of different types of contextual guidance and difficulty levels of the dataset. Our project page is here: https://wikichao.github.io/VisAH/.

📄 PDF Abstract BibTeX arXiv:2505.12154

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An EEG-based Stereoscopic Research to Reveal the Brain's Response to What Happens Before and After Watching 2D and 3D Movies

2019-03-13

Despite knowing the reality of three-dimensional (3D) technology in the form of eye fatigue, this technology continues to be retained by people (especially the young community). To check what happens before and after wat…

BenchmarkingEEGElectroencephalogram (EEG)

Extending MovieLens-32M to Provide New Evaluation Objectives

2025-04-02 · Mark D. Smucker, Houmaan Chamani

Offline evaluation of recommender systems has traditionally treated the problem as a machine learning problem. In the classic case of recommending movies, where the user has provided explicit ratings of which movies they…

Information RetrievalRecommendation Systems

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions

2026-04-17 · Ze Dong, Hao Shi, Zejia Gao, Zhonghua Yi 외 arxiv

Embodied robotic agents often perceive movies through an egocentric screen-view interface rather than native cinematic footage, introducing domain shifts such as viewpoint distortion, scale variation, illumination change…

Domain GeneralizationMultimodal Reasoning

Watching Too Much Television is Good: Self-Supervised Audio-Visual Representation Learning from Movies and TV Shows

2021-06-16 · NeurIPS 2021 12 · Mahdi M. Kalayeh, Nagendra Kamath, Lingyi Liu, Ashok Chandrashekar

The abundance and ease of utilizing sound, along with the fact that auditory clues reveal so much about what happens in the scene, make the audio-visual space a perfectly intuitive choice for self-supervised representati…

Contrastive LearningRepresentation LearningSelf-Supervised Learning

Folksonomication: Predicting Tags for Movies from Plot Synopses Using Emotion Flow Encoded Neural Network

2018-08-15 · COLING 2018 8 · Sudipta Kar, Suraj Maharjan, Thamar Solorio

Folksonomy of movies covers a wide range of heterogeneous information about movies, like the genre, plot structure, visual experiences, soundtracks, metadata, and emotional experiences from watching a movie. Being able t…

Retrieval