Automated Audio Captioning using Transfer Learning and Reconstruction Latent Space Similarity Regularization
In this paper, we examine the use of Transfer Learning using Pretrained Audio Neural Networks (PANNs), and propose an architecture that is able to better leverage the acoustic features provided by PANNs for the Automated Audio Captioning Task. We also introduce a novel self-supervised objective, Reconstruction Latent Space Similarity Regularization (RLSSR). The RLSSR module supplements the training of the model by minimizing the similarity between the encoder and decoder embedding. The combination of both methods allows us to surpass state of the art results by a significant margin on the Clotho dataset across several metrics and benchmarks.
Code (0)
등록된 구현이 없습니다.
Tasks
Audio captioningDecoderTransfer LearningSimilar Papers 제목 키워드 기반
Parameter Efficient Audio Captioning With Faithful Guidance Using Audio-text Shared Latent Representation
There has been significant research on developing pretrained transformer architectures for multimodal-to-text generation tasks. Albeit performance improvements, such models are frequently overparameterized, hence suffer …
Audio captioningData AugmentationHallucinationSemantic Similarity+2Impact of visual assistance for automated audio captioning
We study the impact of visual assistance for automated audio captioning. Utilizing multi-encoder transformer architectures, which have previously been employed to introduce vision-related information in the context of so…
Audio captioningEvent DetectionSound Event DetectionTransfer LearningCL4AC: A Contrastive Loss for Audio Captioning
Automated Audio captioning (AAC) is a cross-modal translation task that aims to use natural language to describe the content of an audio clip. As shown in the submissions received for Task 6 of the DCASE 2021 Challenges,…
Audio captioningDecoderTranslationEfficient Audio Captioning Transformer with Patchout and Text Guidance
Automated audio captioning is multi-modal translation task that aim to generate textual descriptions for a given audio clip. In this paper we propose a full Transformer architecture that utilizes Patchout as proposed in …
Audio captioningCaption GenerationSemantic SimilaritySemantic Textual Similarity+1An Encoder-Decoder Based Audio Captioning System With Transfer and Reinforcement Learning
Automated audio captioning aims to use natural language to describe the content of audio data. This paper presents an audio captioning system with an encoder-decoder architecture, where the decoder predicts words based o…
Audio captioningDecoderreinforcement-learningReinforcement Learning+2