paper-with-me

Papers

Visual and audio scene classification for detecting discrepancies in video: a baseline method and experimental protocol

2024-05-01 · Konstantinos Apostolidis, Jakob Abesser, Luca Cuccovillo, Vasileios Mezaris

This paper presents a baseline approach and an experimental protocol for a specific content verification problem: detecting discrepancies between the audio and video modalities in multimedia content. We first design and optimize an audio-visual scene classifier, to compare with existing classification baselines that use both modalities. Then, by applying this classifier separately to the audio and the visual modality, we can detect scene-class inconsistencies between them. To facilitate further research and provide a common evaluation platform, we introduce an experimental protocol and a benchmark dataset simulating such inconsistencies. Our approach achieves state-of-the-art results in scene classification and promising outcomes in audio-visual discrepancies detection, highlighting its potential in content verification applications.

📄 PDF Abstract BibTeX arXiv:2405.00384

Code (1)

idt-iti/visual-audio-discrepancy-detection 공식 구현 pytorch

Tasks

Scene Classification

Similar Papers 제목 키워드 기반

Echo-Reconstruction: Audio-Augmented 3D Scene Reconstruction

2021-10-05 · Justin Wilson, Nicholas Rewkowski, Ming C. Lin, Henry Fuchs

Reflective and textureless surfaces such as windows, mirrors, and walls can be a challenge for object and scene reconstruction. These surfaces are often poorly reconstructed and filled with depth discontinuities and hole…

3D Reconstruction3D Scene ReconstructionClassificationDepth Estimation+1

An Audio-Visual Dataset and Deep Learning Frameworks for Crowded Scene Classification

2021-12-16 · Lam Pham, Dat Ngo, Phu X. Nguyen, Truong Hoang 외

This paper presents a task of audio-visual scene classification (SC) where input videos are classified into one of five real-life crowded scenes: 'Riot', 'Noise-Street', 'Firework-Event', 'Music-Event', and 'Sport-Atmosp…

Deep LearningScene Classification

Discrepancy-Aware Attention Network for Enhanced Audio-Visual Zero-Shot Learning

2024-12-16 · RunLin Yu, Yipu Gong, Wenrui Li, Aiwen Sun 외

Audio-visual Zero-Shot Learning (ZSL) has attracted significant attention for its ability to identify unseen classes and perform well in video classification tasks. However, modal imbalance in (G)ZSL leads to over-relian…

Video ClassificationZero-Shot Learning

A proto-object based audiovisual saliency map

2020-03-15 · Sudarshan Ramenahalli

Natural environment and our interaction with it is essentially multisensory, where we may deploy visual, tactile and/or auditory senses to perceive, learn and interact with our environment. Our objective in this study is…

ObjectvalidVideo Compression

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning

2026-06-10 · Zihan Zhang, Jie Hong, Siyuan Fan, Yanghao Zhou 외 arxiv

Audio-visual Generalized Zero-shot Learning (AV-GZSL) is a challenging task that aims to classify both seen and unseen objects or scenes by integrating data from audio and visual modalities. Recent studies primarily focu…

Generalized Zero-Shot Learning