paper-with-me

Papers

CoLLAT: On Adding Fine-grained Audio Understanding to Language Models using Token-Level Locked-Language Tuning

2023-09-21 · NeurIPS 2023 11

Humans can easily understand various audio concepts, but conventional audio classification models fail due to their inability to predict unseen classes during training. To address this challenge, recent literature has explored contrastive language-audio pretraining to learn an audio understanding model using natural language supervision from a pretrained language model. However, despite their reasonable zero-shot performance in audio understanding, these models typically fail to achieve optimal performance while preserving the text understanding capabilities of the pretrained language model. They also perform poorly when comprehending audio clips with multiple audio concepts. To bridge these gaps, we propose $CoLLAT$: $Co$ntrastive $L$ocked $L$anguage and $A$udio $T$uning. This is a framework to effectively learn an audio understanding model with a locked language model, which is learned using a novel pretraining objective for audio-to-text grounding to yield fine-grained audio understanding. Our extensive experiments, which include several downstream applications such as audio classification, cross-modal retrieval, and audio-guided image generation, demonstrate that $CoLLAT$ yields state-of-the-art performance for audio understanding. Additionally, it unlocks audio guidance to applications built on top of pretrained language models.Submission Number: 13079

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning

2026-07-22 · Siqian Tong, Xuan Li, Chaozhuo Li, Baolong Bi 외 arxiv

Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-t…

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks

2025-07-31 · Yadong Niu, Tianzi Wang, Heinrich Dinkel, Xingwei Sun 외 arxiv

While large audio-language models have advanced open-ended audio understanding, they still fall short of nuanced human-level comprehension. This gap persists largely because current benchmarks, limited by data annotation…

Semantic Similarity

RA-SSU: Towards Fine-Grained Audio-Visual Learning with Region-Aware Sound Source Understanding

2026-03-10 · Muyi Sun, Yixuan Wang, Hong Wang, Chen Su 외 arxiv

Audio-Visual Learning (AVL) is one fundamental task of multi-modality learning and embodied intelligence, displaying the vital role in scene understanding and interaction. However, previous researchers mostly focus on ex…

audio-visual event localizationSound Source LocalizationScene Understanding

Towards Universal Video MLLMs with Attribute-Structured and Quality-Verified Instructions

2026-02-13 · Yunheng Li, Hengrui Zhang, Meng-Hao Guo, Wenzhao Gao 외 arxiv

Universal video understanding requires modeling fine-grained visual and audio information over time in diverse real-world scenarios. However, the performance of existing models is primarily constrained by video-instructi…

Instruction Following

Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt

2026-04-15 · Yanfeng Shi, Pengfei Cai, Jun Liu, Qing Gu 외 arxiv

Large Audio-Language Models (LALMs) enable general audio understanding and demonstrate remarkable performance across various audio tasks. However, these models still face challenges in temporal perception (e.g., inferrin…

Reinforcement LearningSound Event DetectionAudio captioning