Audio captioning
2개 벤치마크 · 논문 140편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Clotho: An Audio Captioning Dataset
WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research
M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
Papers
MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on gener…
Audio captioningAn Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations
Recent advancements in automated audio captioning (AAC) have shifted from monolithic sentence generation toward structured formats that explicitly disentangle distinct acoustic and semantic properties. However, evaluatin…
Logical ReasoningAudio captioningCARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning
Modern automated audio captioning systems pair a frozen audio encoder with a large language model (LLM) via a trainable projector, incurring the encoder's inference cost and bottlenecking the model through its fixed acou…
Audio captioningMOSS-Audio Technical Report
MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understanding, supporting audio captioning, time-aware question answering, timestamped transcription, and audio-grounded reasoning. …
Instruction FollowingQuestion AnsweringAudio captioningLocality Matters for Training-Free Audio Token Compression in Audio-Language Models
Audio-language models (ALMs) are increasingly used for audio captioning, question answering, and open-ended audio understanding, but their inference cost remains high when audio inputs are represented as long prefix-toke…
Question AnsweringAudio captioningTowards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt
Large Audio-Language Models (LALMs) enable general audio understanding and demonstrate remarkable performance across various audio tasks. However, these models still face challenges in temporal perception (e.g., inferrin…
Reinforcement LearningSound Event DetectionAudio captioning