paper-with-me

Papers

Do Audio-Language Models Understand Linguistic Variations?

2024-10-21 · Ramaneswaran Selvakumar, Sonal Kumar, Hemant Kumar Giri, Nishit Anand, Ashish Seth, Sreyan Ghosh, Dinesh Manocha

Open-vocabulary audio language models (ALMs), like Contrastive Language Audio Pretraining (CLAP), represent a promising new paradigm for audio-text retrieval using natural language queries. In this paper, for the first time, we perform controlled experiments on various benchmarks to show that existing ALMs struggle to generalize to linguistic variations in textual queries. To address this issue, we propose RobustCLAP, a novel and compute-efficient technique to learn audio-language representations agnostic to linguistic variations. Specifically, we reformulate the contrastive loss used in CLAP architectures by introducing a multi-view contrastive learning objective, where paraphrases are treated as different views of the same audio scene and use this for training. Our proposed approach improves the text-to-audio retrieval performance of CLAP by 0.8%-13% across benchmarks and enhances robustness to linguistic variation.

📄 PDF Abstract BibTeX arXiv:2410.16505

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningNatural Language QueriesRetrievalText RetrievalText to Audio Retrieval

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox

2026-05-26 · Jiacheng Pang, Ashutosh Chaubey, Mohammad Soleymani arxiv

Audio large language models (Audio LLMs) demonstrate strong performance on speech understanding tasks, yet their ability to understand paralinguistic information remains limited. To systematically quantify this issue, we…

Speech Synthesis

Eureka-Audio: Triggering Audio Intelligence in Compact Language Models

2026-02-15 · Dan Zhang, Yishu Lei, Jing Hu, Shuwei He 외 arxiv

We present Eureka-Audio, a compact yet high-performance audio language model that achieves competitive performance against models that are 4 to 18 times larger across a broad range of audio understanding benchmarks. Desp…

Instruction FollowingSpeech RecognitionAudio captioning

Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations

2025-10-19 · Bo-Han Feng, Chien-Feng Liu, Yu-Hsuan Li Liang, Chih-Kai Yang 외 arxiv

Large audio-language models (LALMs) extend text-based LLMs with auditory understanding, offering new opportunities for multimodal applications. While their perception, reasoning, and task performance have been widely stu…

Audio-Linguistic Embeddings for Spoken Sentences

2019-02-20 · Albert Haque, Michelle Guo, Prateek Verma, Li Fei-Fei

We propose spoken sentence embeddings which capture both acoustic and linguistic content. While existing works operate at the character, phoneme, or word level, our method learns long-term dependencies by modeling speech…

DecoderEmotion RecognitionSentenceSentence Embeddings+3

Multilingual and Multi-Accent Jailbreaking of Audio LLMs

2025-04-01 · Jaechul Roh, Virat Shejwalkar, Amir Houmansadr

Large Audio Language Models (LALMs) have significantly advanced audio understanding but introduce critical security risks, particularly through audio jailbreaks. While prior work has focused on English-centric attacks, w…