paper-with-me

홈 › Papers

Investigations in Audio Captioning: Addressing Vocabulary Imbalance and Evaluating Suitability of Language-Centric Performance Metrics

2022-11-12 · Sandeep Kothinti, Dimitra Emmanouilidou

The analysis, processing, and extraction of meaningful information from sounds all around us is the subject of the broader area of audio analytics. Audio captioning is a recent addition to the domain of audio analytics, a cross-modal translation task that focuses on generating natural descriptions from sound events occurring in an audio stream. In this work, we identify and improve on three main challenges in automated audio captioning: i) data scarcity, ii) imbalance or limitations in the audio captions vocabulary, and iii) the proper performance evaluation metric that can best capture both auditory and semantic characteristics. We find that generally adopted loss functions can result in an unfair vocabulary imbalance during model training. We propose two audio captioning augmentation methods that enrich the training dataset and the vocabulary size. We further underline the need for in-domain pretraining by exploring the suitability of audio encoders that were previously trained on different audio tasks. Finally, we systematically explore five performance metrics borrowed from the image captioning domain and highlight their limitations for the audio domain.

📄 PDF Abstract BibTeX arXiv:2211.06547

Code (0)

등록된 구현이 없습니다.

Tasks

Audio captioningImage Captioning

Similar Papers 제목 키워드 기반

Multi-task Regularization Based on Infrequent Classes for Audio Captioning

2020-07-09 · Emre Çakır, Konstantinos Drossos, Tuomas Virtanen

Audio captioning is a multi-modal task, focusing on using natural language for describing the contents of general audio. Most audio captioning methods are based on deep neural networks, employing an encoder-decoder schem…

Audio captioningDecoder

ChordFormer: A Conformer-Based Architecture for Large-Vocabulary Audio Chord Recognition

2025-02-17 · Muhammad Waseem Akram, Stefano Dettori, Valentina Colla, Giorgio Carlo Buttazzo

Chord recognition serves as a critical task in music information retrieval due to the abstract and descriptive nature of chords in music analysis. While audio chord recognition systems have achieved significant accuracy …

Chord RecognitionDescriptiveInformation RetrievalMusic Information Retrieval

Diversity and bias in audio captioning datasets

2022-11-15 · DCASE workshop 2022 11 · Irene Martin-Morato, Annamaria Mesaros

Describing soundscapes in sentences allows better understand- ing of the acoustic scene than a single label indicating the acoustic scene class or a set of audio tags indicating the sound events active in the audio clip.…

Audio captioningDiversity

Mitigating Open-Vocabulary Caption Hallucinations

2023-12-06 · Assaf Ben-Kish, Moran Yanuka, Morris Alper, Raja Giryes 외

While recent years have seen rapid progress in image-conditioned text generation, image captioning still suffers from the fundamental issue of hallucinations, namely, the generation of spurious details that cannot be inf…

DiversityHallucinationImage CaptioningObject+2

Classifier-Guided Captioning Across Modalities

2025-01-03 · Ariel Shaulov, Tal Shaharabany, Eitan Shaar, Gal Chechik 외

Most current captioning systems use language models trained on data from specific settings, such as image-based captioning via Amazon Mechanical Turk, limiting their ability to generalize to other modality distributions …

Audio captioningVideo CaptioningZero-shot Audio Captioning