paper-with-me

홈 › Papers

The NTT DCASE2020 Challenge Task 6 system: Automated Audio Captioning with Keywords and Sentence Length Estimation

2020-07-01 · Yuma Koizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, Kunio Kashino

This technical report describes the system participating to the Detection and Classification of Acoustic Scenes and Events (DCASE) 2020 Challenge, Task 6: automated audio captioning. Our submission focuses on solving two indeterminacy problems in automated audio captioning: word selection indeterminacy and sentence length indeterminacy. We simultaneously solve the main caption generation and sub indeterminacy problems by estimating keywords and sentence length through multi-task learning. We tested a simplified model of our submission using the development-testing dataset. Our model achieved 20.7 SPIDEr score where that of the baseline system was 5.4.

📄 PDF Abstract BibTeX arXiv:2007.00225

Code (0)

등록된 구현이 없습니다.

Tasks

Audio captioningCaption GenerationMulti-Task LearningSentence

Similar Papers 제목 키워드 기반

THE DCASE 2021 CHALLENGE TASK 6 SYSTEM: AUTOMATED AUDIO CAPTIONING WITH WEAKLY SUPERVISED PRE-TRAING AND WORD SELECTION METHODS

2021-07-06 · DCASE workshop 2021 7 · Weiqiang Yuan ∗, Qichen Han∗, Dong Liu, Xiang Li 외

This technical report describes the system participating to the De- tection and Classification of Acoustic Scenes and Events (DCASE) 2021 Challenge, Task 6: automated audio captioning. We use encoder-decoder modeling …

Audio captioningCaption GenerationDecoder

Expanding on EnCLAP with Auxiliary Retrieval Model for Automated Audio Captioning

2024-09-02 · Jaeyeon Kim, JaeYoon Jung, Minjeong Jeon, Sang Hoon Woo 외

In this technical report, we describe our submission to DCASE2024 Challenge Task6 (Automated Audio Captioning) and Task8 (Language-based Audio Retrieval). We develop our approach building upon the EnCLAP audio captioning…

Audio captioningRerankingRetrieval

Language-based Audio Retrieval Task in DCASE 2022 Challenge

2022-06-13 · Huang Xie, Samuel Lipping, Tuomas Virtanen

Language-based audio retrieval is a task, where natural language textual captions are used as queries to retrieve audio signals from a dataset. It has been first introduced into DCASE 2022 Challenge as Subtask 6B of task…

Audio captioningRetrieval

Language-based Audio Retrieval Task in DCASE 2022 Challenge

2022-09-20 · Huang Xie, Samuel Lipping, Tuomas Virtanen

Language-based audio retrieval is a task, where natural language textual captions are used as queries to retrieve audio signals from a dataset. It has been first introduced into DCASE 2022 Challenge as Subtask 6B of task…

Audio captioningRetrieval

General audio tagging with ensembling convolutional neural network and statistical features

2018-10-30 · Kele Xu, Boqing Zhu, Qiuqiang Kong, Haibo Mi 외

Audio tagging aims to infer descriptive labels from audio clips. Audio tagging is challenging due to the limited size of data and noisy labels. In this paper, we describe our solution for the DCASE 2018 Task 2 general au…

Audio TaggingDescriptiveEnsemble LearningTask 2