paper-with-me

Audio captioning

2개 벤치마크 · 논문 140편 · 이 태스크의 논문 보기 →

Benchmarks

AudioCaps

결과 18개

Clotho

결과 11개

Most implemented

Clotho: An Audio Captioning Dataset

2019-10-21 · 구현 7개

Papers

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

2026-07-29 · Weijie Wu, Junbo Li, Lin Li, Jun Fang 외 arxiv

With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on gener…

Audio captioning

An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations

2026-07-23 · Liang-Yuan Wu, Sripathi Sridhar, Mark Cartwright, Magdalena Fuentes arxiv

Recent advancements in automated audio captioning (AAC) have shifted from monolithic sentence generation toward structured formats that explicitly disentangle distinct acoustic and semantic properties. However, evaluatin…

Logical ReasoningAudio captioning

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning

2026-07-06 · Ganesh Pavan Kartikeya Bharadwaj Kolluri, Yuchen Zhang, Michael Kampouridis, Ravi Shekhar arxiv

Modern automated audio captioning systems pair a frozen audio encoder with a large language model (LLM) via a trainable projector, incurring the encoder's inference cost and bottlenecking the model through its fixed acou…

Audio captioning

MOSS-Audio Technical Report

2026-06-01 · Chen Yang, Chufan Yu, Hanfu Chen, Jie Zhu 외 arxiv

MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understanding, supporting audio captioning, time-aware question answering, timestamped transcription, and audio-grounded reasoning. …

Instruction FollowingQuestion AnsweringAudio captioning

Locality Matters for Training-Free Audio Token Compression in Audio-Language Models

2026-05-24 · Jiale Luo, Xiaoyu Liang, Haoji Hu arxiv

Audio-language models (ALMs) are increasingly used for audio captioning, question answering, and open-ended audio understanding, but their inference cost remains high when audio inputs are represented as long prefix-toke…

Question AnsweringAudio captioning

Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt

2026-04-15 · Yanfeng Shi, Pengfei Cai, Jun Liu, Qing Gu 외 arxiv

Large Audio-Language Models (LALMs) enable general audio understanding and demonstrate remarkable performance across various audio tasks. However, these models still face challenges in temporal perception (e.g., inferrin…

Reinforcement LearningSound Event DetectionAudio captioning

전체 140편 보기 →