paper-with-me

Papers

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples

2025-05-20 · Chun-Yi Kuan, Hung-Yi Lee

Recent advancements in audio-aware large language models (ALLMs) enable them to process and understand audio inputs. However, these models often hallucinate non-existent sound events, reducing their reliability in real-world applications. To address this, we propose LISTEN (Learning to Identify Sounds Through Extended Negative Samples), a contrastive-like training method that enhances ALLMs' ability to distinguish between present and absent sounds using synthesized data from the backbone LLM. Unlike prior approaches, our method requires no modification to LLM parameters and efficiently integrates audio representations via a lightweight adapter. Experiments show that LISTEN effectively mitigates hallucinations while maintaining impressive performance on existing audio question and reasoning benchmarks. At the same time, it is more efficient in both data and computation.

📄 PDF Abstract BibTeX arXiv:2505.14518

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Teaching Physical Awareness to LLMs through Sounds

2025-06-10 · Weiguo Wang, Andy Nie, Wenrui Zhou, Yi Kai 외

Large Language Models (LLMs) have shown remarkable capabilities in text and multimodal processing, yet they fundamentally lack physical awareness--understanding of real-world physical phenomena. In this work, we present …

Direction of Arrival Estimation

ChronusOmni: Improving Time Awareness of Omni Large Language Models

2025-12-10 · Yijing Chen, Yihan Wu, Kaisi Guan, Yuchen Ren 외 arxiv

Time awareness is a fundamental ability of omni large language models, especially for understanding long videos and answering complex questions. Previous approaches mainly target vision-language scenarios and focus on th…

Reinforcement Learning

Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding

2025-06-08 · Tzu-wen Hsu, Ke-Han Lu, Cheng-Han Chiang, Hung-Yi Lee

Large Audio-Language Models (LALMs) can take audio and text as the inputs and answer questions about the audio. While prior LALMs have shown strong performance on standard benchmarks, there has been alarming evidence tha…

HallucinationObject Hallucination

Multimodal Classification of Teaching Activities from University Lecture Recordings

2023-12-24 · Oscar Sapena, Eva Onaindia

The way of understanding online higher education has greatly changed due to the worldwide pandemic situation. Teaching is undertaken remotely, and the faculty incorporate lecture audio recordings as part of the teaching …

ClassificationLanguage Modelling

Learning Spatially-Aware Language and Audio Embeddings

2024-09-17 · Bhavika Devnani, Skyler Seto, Zakaria Aldeneh, Alessandro Toso 외

Humans can picture a sound scene given an imprecise natural language description. For example, it is easy to imagine an acoustic environment given a phrase like "the lion roar came from right behind me!". For a machine t…

AttributeContrastive LearningPositionSemantic Retrieval+1