Effects of Word-frequency based Pre- and Post- Processings for Audio Captioning
The system we used for Task 6 (Automated Audio Captioning)of the Detection and Classification of Acoustic Scenes and Events(DCASE) 2020 Challenge combines three elements, namely, dataaugmentation, multi-task learning, and post-processing, for audiocaptioning. The system received the highest evaluation scores, butwhich of the individual elements most fully contributed to its perfor-mance has not yet been clarified. Here, to asses their contributions,we first conducted an element-wise ablation study on our systemto estimate to what extent each element is effective. We then con-ducted a detailed module-wise ablation study to further clarify thekey processing modules for improving accuracy. The results showthat data augmentation and post-processing significantly improvethe score in our system. In particular, mix-up data augmentationand beam search in post-processing improve SPIDEr by 0.8 and 1.6points, respectively.
Code (0)
등록된 구현이 없습니다.
Tasks
Audio captioningData AugmentationMulti-Task LearningSimilar Papers 제목 키워드 기반
Inference-time Scaling for Diffusion-based Audio Super-resolution
Diffusion models have demonstrated remarkable success in generative tasks, including audio super-resolution (SR). In many applications like movie post-production and album mastering, substantial computational budgets are…
Audio Super-ResolutionPost-Processing Independent Evaluation of Sound Event Detection Systems
Due to the high variation in the application requirements of sound event detection (SED) systems, it is not sufficient to evaluate systems only in a single operating mode. Therefore, the community recently adopted the po…
Event DetectionSound Event DetectionFrequency effects in Linear Discriminative Learning
Word frequency is a strong predictor in most lexical processing tasks. Thus, any model of word recognition needs to account for how word frequency effects arise. The Discriminative Lexicon Model (DLM; Baayen et al., 2018…
Incremental LearningA large-scale study of the effects of word frequency and predictability in naturalistic reading
A number of psycholinguistic studies have factorially manipulated words{'} contextual predictabilities and corpus frequencies and shown separable effects of each on measures of human sentence processing, a pattern which …
RetrievalSentenceModulation Extraction for LFO-driven Audio Effects
Low frequency oscillator (LFO) driven audio effects such as phaser, flanger, and chorus, modify an input signal using time-varying filters and delays, resulting in characteristic sweeping or widening effects. It has been…