paper-with-me

Papers

What Do Language Models Hear? Probing for Auditory Representations in Language Models

2024-02-26 · Jerry Ngo, Yoon Kim

This work explores whether language models encode meaningfully grounded representations of sounds of objects. We learn a linear probe that retrieves the correct text representation of an object given a snippet of audio related to that object, where the sound representation is given by a pretrained audio model. This probe is trained via a contrastive loss that pushes the language representations and sound representations of an object to be close to one another. After training, the probe is tested on its ability to generalize to objects that were not seen during training. Across different language models and audio models, we find that the probe generalization is above chance in many cases, indicating that despite being trained only on raw text, language models encode grounded knowledge of sounds for some objects.

📄 PDF Abstract BibTeX arXiv:2402.16998

Code (0)

등록된 구현이 없습니다.

Tasks

Object

Similar Papers 제목 키워드 기반

Eidos: An Open-Source Auditory Periphery Modeling Toolkit and Evaluation of Cross-Lingual Phonemic Contrasts

2020-05-01 · LREC 2020 5 · Alex Gutkin, er

Many analytical models that mimic, in varying degree of detail, the basic auditory processes involved in human hearing have been developed over the past decades. While the auditory periphery mechanisms responsible for tr…

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning

2026-07-22 · Siqian Tong, Xuan Li, Chaozhuo Li, Baolong Bi 외 arxiv

Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-t…

Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models

2026-07-11 · Yun-Shao Tsai, Chun-Wei Chen, Chee-En Yu, Yi-Cheng Lin 외 arxiv

Sound symbolism, the human tendency to map speech sounds to perceptual qualities such as roundness or sharpness, arises primarily from the acoustics of speech rather than spelling. Whether Speech Language Models (SLMs) s…

What does a network layer hear? Analyzing hidden representations of end-to-end ASR through speech synthesis

2019-11-04 · Chung-Yi Li, Pei-Chieh Yuan, Hung-Yi Lee

End-to-end speech recognition systems have achieved competitive results compared to traditional systems. However, the complex transformations involved between layers given highly variable acoustic signals are hard to ana…

Speaker VerificationSpeech Enhancementspeech-recognitionSpeech Recognition+1

AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?

2025-09-22 · Hyunjong Ok, Suho Yoo, Hyeonjun Kim, Jaeho Lee arxiv

Even without directly hearing sounds, humans can effortlessly reason about auditory properties, such as pitch, loudness, or sound-source associations, drawing on auditory commonsense. In contrast, language models often l…