paper-with-me

홈 › Papers

Imagine to Hear: Auditory Knowledge Generation can be an Effective Assistant for Language Models

2025-03-21 · Suho Yoo, Hyunjong Ok, Jaeho Lee

Language models pretrained on text-only corpora often struggle with tasks that require auditory commonsense knowledge. Previous work addresses this problem by augmenting the language model to retrieve knowledge from external audio databases. This approach has several limitations, such as the potential lack of relevant audio in databases and the high costs associated with constructing the databases. To address these issues, we propose Imagine to Hear, a novel approach that dynamically generates auditory knowledge using generative models. Our framework detects multiple audio-related textual spans from the given prompt and generates corresponding auditory knowledge. We develop several mechanisms to efficiently process multiple auditory knowledge, including a CLAP-based rejection sampler and a language-audio fusion module. Our experiments show that our method achieves state-of-the-art performance on AuditoryBench without relying on external databases, highlighting the effectiveness of our generation-based approach.

📄 PDF Abstract BibTeX arXiv:2503.16853

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Decoding Imagined Auditory Pitch Phenomena with an Autoencoder Based Temporal Convolutional Architecture

2023-05-15 · Sean Paulsen, Lloyd May, Michael Casey

Stimulus decoding of functional Magnetic Resonance Imaging (fMRI) data with machine learning models has provided new insights about neural representational spaces and task-related dynamics. However, the scarcity of label…

AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?

2025-09-22 · Hyunjong Ok, Suho Yoo, Hyeonjun Kim, Jaeho Lee arxiv

Even without directly hearing sounds, humans can effortlessly reason about auditory properties, such as pitch, loudness, or sound-source associations, drawing on auditory commonsense. In contrast, language models often l…

Using Neurogram Similarity Index Measure (NSIM) to Model Hearing Loss and Cochlear Neural Degeneration

2025-06-15 · Ahsan J. Cheema, Sunil Puria

Trouble hearing in noisy situations remains a common complaint for both individuals with hearing loss and individuals with normal hearing. This is hypothesized to arise due to condition called: cochlear neural degenerati…

Phoneme Recognition

Hearing-Loss Compensation Using Deep Neural Networks: A Framework and Results From a Listening Test

2024-03-15 · Peter Leer, Jesper Jensen, Laurel H. Carney, Zheng-Hua Tan 외

This article investigates the use of deep neural networks (DNNs) for hearing-loss compensation. Hearing loss is a prevalent issue affecting millions of people worldwide, and conventional hearing aids have limitations in …

Music ClassificationSpeaker Identificationspeech-recognitionSpeech Recognition

Spatial Speech Translation: Translating Across Space With Binaural Hearables

2025-04-25 · Tuochao Chen, Qirui Wang, Runlin He, Shyam Gollakota

Imagine being in a crowded space where people speak a different language and having hearables that transform the auditory space into your native language, while preserving the spatial cues for all speakers. We introduce …

blind source separationTranslation