paper-with-me

Papers

The Influence of Audio on Video Memorability with an Audio Gestalt Regulated Video Memorability System

2021-04-23 · Lorin Sweeney, Graham Healy, Alan F. Smeaton

Memories are the tethering threads that tie us to the world, and memorability is the measure of their tensile strength. The threads of memory are spun from fibres of many modalities, obscuring the contribution of a single fibre to a thread's overall tensile strength. Unfurling these fibres is the key to understanding the nature of their interaction, and how we can ultimately create more meaningful media content. In this paper, we examine the influence of audio on video recognition memorability, finding evidence to suggest that it can facilitate overall video recognition memorability rich in high-level (gestalt) audio features. We introduce a novel multimodal deep learning-based late-fusion system that uses audio gestalt to estimate the influence of a given video's audio on its overall short-term recognition memorability, and selectively leverages audio features to make a prediction accordingly. We benchmark our audio gestalt based system on the Memento10k short-term video memorability dataset, achieving top-2 state-of-the-art results.

📄 PDF Abstract BibTeX arXiv:2104.11568

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Deep LearningVideo Recognition

Similar Papers 제목 키워드 기반

Leveraging Audio Gestalt to Predict Media Memorability

2020-12-31 · Lorin Sweeney, Graham Healy, Alan F. Smeaton

Memorability determines what evanesces into emptiness, and what worms its way into the deepest furrows of our minds. It is the key to curating more meaningful media content as we wade through daily digital torrents. The …

Multimodal Deep Learning

Multi-modal Ensemble Models for Predicting Video Memorability

2021-02-01 · Tony Zhao, Irving Fang, Jeffrey Kim, Gerald Friedland

Modeling media memorability has been a consistent challenge in the field of machine learning. The Predicting Media Memorability task in MediaEval2020 is the latest benchmark among similar challenges addressing this topic…

BIG-bench Machine Learning

Coherent Audio-Visual Editing via Conditional Audio Generation Following Video Edits

2025-12-08 · Masato Ishii, Akio Hayakawa, Takashi Shibuya, Yuki Mitsufuji arxiv

We introduce a novel pipeline for joint audio-visual editing that enhances the coherence between edited video and its accompanying audio. Our approach first applies state-of-the-art video editing techniques to produce th…

Data AugmentationAudio Generation

Predicted Cortex Is Not a Domain-General Prior: A Matched-Control Audit of Brain-Encoding Features for Video Memorability

2026-07-12 · Carson Rodrigues arxiv

Brain-encoding foundation models predict fMRI responses to video, audio and text well enough to win the Algonauts 2025 challenge. We ask whether their predicted responses, obtained with no scanner, are a useful feature l…

"Did You Hear That?" Learning to Play Video Games from Audio Cues

2019-06-10 · Raluca D. Gaina, Matthew Stephenson

Game-playing AI research has focused for a long time on learning to play video games from visual input or symbolic information. However, humans benefit from a wider array of sensors which we utilise in order to navigate …

Game DesignNavigateQ-Learning