paper-with-me

Papers Zero-shot Audio Captioning

“Zero-shot Audio Captioning” 태그가 달린 논문 8편 · 필터 해제

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings

2026-05-28 · Yonggang Zhu, Liting Gao, Aidong Men, Wenwu Wang arxiv

Contrastive Language-Audio Pretraining (CLAP) models are widely used for audio understanding and support modality-agnostic condition swapping in many zero-shot applications. However, their performance is heavily affected…

Zero-shot Audio CaptioningDimensionality Reduction

MAGIC-Enhanced Keyword Prompting for Zero-Shot Audio Captioning with CLIP Models

2025-09-16 · Vijay Govindarajan, Pratik Patel, Sahil Tripathi, Md Azizul Hoque 외 arxiv

Automated Audio Captioning (AAC) generates captions for audio clips but faces challenges due to limited datasets compared to image captioning. To overcome this, we propose the zero-shot AAC system that leverages pre-trai…

Zero-shot Audio CaptioningImage Captioning

Classifier-Guided Captioning Across Modalities

2025-01-03 · Ariel Shaulov, Tal Shaharabany, Eitan Shaar, Gal Chechik 외

Most current captioning systems use language models trained on data from specific settings, such as image-based captioning via Amazon Mechanical Turk, limiting their ability to generalize to other modality distributions …

Audio captioningVideo CaptioningZero-shot Audio Captioning

DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning

2024-10-12 · Xiquan Li, Wenxi Chen, Ziyang Ma, Xuenan Xu 외

While automated audio captioning (AAC) has made notable progress, traditional fully supervised AAC models still face two critical challenges: the need for expensive audio-text pair data for training and performance degra…

Audio captioningLarge Language ModelRetrievalRetrieval-augmented Generation+1

An Eye for an Ear: Zero-shot Audio Description Leveraging an Image Captioner using Audiovisual Distribution Alignment

2024-10-08 · Hugo Malard, Michel Olvera, Stéphane Lathuiliere, Slim Essid

Multimodal large language models have fueled progress in image captioning. These models, fine-tuned on vast image datasets, exhibit a deep understanding of semantic concepts. In this work, we show that this ability can b…

Audio captioningContrastive LearningImage CaptioningRepresentation Learning+1

Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

2024-02-02 · Zhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping 외

Augmenting large language models (LLMs) to understand audio -- including non-speech sounds and non-verbal speech -- is critically important for diverse real-world applications of LLMs. In this paper, we propose Audio Fla…

Acoustic Scene ClassificationAudio captioningFew-Shot LearningIn-Context Learning+5

Zero-shot audio captioning with audio-language model guidance and audio context keywords

2023-11-14 · Leonard Salewski, Stefan Fauth, A. Sophia Koepke, Zeynep Akata

Zero-shot audio captioning aims at automatically generating descriptive textual captions for audio content without prior training for this task. Different from speech recognition which translates audio content that conta…

Audio captioningDescriptiveImage CaptioningLanguage Modeling+5

Zero-Shot Audio Captioning via Audibility Guidance

2023-09-07 · Tal Shaharabany, Ariel Shaulov, Lior Wolf

The task of audio captioning is similar in essence to tasks such as image and video captioning. However, it has received much less attention. We propose three desiderata for captioning audio -- (i) fluency of the generat…

Language ModelingZero-shot Audio Captioning
1–8 / 8