paper-with-me

AudioCaps

홈페이지 · 논문 279편

AudioCaps is a dataset of sounds with event descriptions that was introduced for the task of audio captioning, with sounds sourced from the AudioSet dataset. Annotators were provided the audio tracks together with category hints (and with additional video hints if needed). Source: [Audio Retrieval with Natural Language Queries](/paper/audio-retrieval-with-natural-language-queries) Image source: https://audiocaps.github.io/

TextsAudio

벤치마크

Audio Generation on AudioCaps 결과 23개
Audio captioning on AudioCaps 결과 18개
Text to Audio Retrieval on AudioCaps 결과 11개
Retrieval-augmented Few-shot In-context Audio Captioning on AudioCaps 결과 5개
Zero-shot Audio Captioning on AudioCaps 결과 4개
Target Sound Extraction on AudioCaps 결과 3개