Audio-based Near-Duplicate Video Retrieval with Audio Similarity Learning
In this work, we address the problem of audio-based near-duplicate video retrieval. We propose the Audio Similarity Learning (AuSiL) approach that effectively captures temporal patterns of audio similarity between video pairs. For the robust similarity calculation between two videos, we first extract representative audio-based video descriptors by leveraging transfer learning based on a Convolutional Neural Network (CNN) trained on a large scale dataset of audio events, and then we calculate the similarity matrix derived from the pairwise similarity of these descriptors. The similarity matrix is subsequently fed to a CNN network that captures the temporal structures existing within its content. We train our network following a triplet generation process and optimizing the triplet loss function. To evaluate the effectiveness of the proposed approach, we have manually annotated two publicly available video datasets based on the audio duplicity between their videos. The proposed approach achieves very competitive results compared to three state-of-the-art methods. Also, unlike the competing methods, it is very robust to the retrieval of audio duplicates generated with speed transformations.
Code (1)
Tasks
RetrievalTransfer LearningTripletVideo RetrievalMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
TimberAgent: Gram-Guided Retrieval for Executable Music Effect Control
Digital audio workstations expose rich effect chains, yet a semantic gap remains between perceptual user intent and low-level signal-processing parameters. We study retrieval-grounded audio effect control, where the outp…
Refining Knowledge Transfer on Audio-Image Temporal Agreement for Audio-Text Cross Retrieval
The aim of this research is to refine knowledge transfer on audio-image temporal agreement for audio-text cross retrieval. To address the limited availability of paired non-speech audio-text data, learning methods for tr…
Image RetrievalRetrievalText RetrievalTransfer LearningAudio-Enhanced Text-to-Video Retrieval using Text-Conditioned Feature Alignment
Text-to-video retrieval systems have recently made significant progress by utilizing pre-trained models trained on large-scale image-text pairs. However, most of the latest methods primarily focus on the video modality w…
RetrievalText to Video RetrievalVideo AlignmentVideo RetrievalCNN Retrieval based Unsupervised Metric Learning for Near-Duplicated Video Retrieval
As important data carriers, the drastically increasing number of multimedia videos often brings many duplicate and near-duplicate videos in the top results of search. Near-duplicate video retrieval (NDVR) can cluster and…
Metric LearningRe-RankingRetrievalVideo RetrievalGeneration or Replication: Auscultating Audio Latent Diffusion Models
The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how we work with audio. In this work, we mak…
AudioCapsMemorizationRetrieval