paper-with-me

홈 › Papers

Seeing isn't Hearing: Benchmarking Vision Language Models at Interpreting Spectrograms

2025-11-17 · Tyler Loakman, Joseph James, Chenghua Lin arxiv

With the rise of Large Language Models (LLMs) and their vision-enabled counterparts (VLMs), numerous works have investigated their capabilities in tasks that fuse the modalities of vision and language. In this work, we benchmark the extent to which VLMs are able to act as highly-trained phoneticians, interpreting spectrograms and waveforms of speech. To do this, we synthesise a novel dataset containing 4k+ English words spoken in isolation alongside stylistically consistent spectrogram and waveform figures. We test the ability of VLMs to understand these representations of speech through a multiple-choice task whereby models must predict the correct phonemic or graphemic transcription of a spoken word when presented amongst 3 distractor transcriptions that have been selected based on their phonemic edit distance to the ground truth. We observe that both zero-shot and finetuned models rarely perform above chance, demonstrating the requirement for specific parametric knowledge of how to interpret such figures, rather than paired samples alone.

📄 PDF Abstract BibTeX arXiv:2511.13225

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound

2025-11-15 · Dengming Zhang, Weitao You, Jingxiong Li, Weishen Lin 외 arxiv

Emotion understanding is critical for making Large Language Models (LLMs) more general, reliable, and aligned with humans. Art conveys emotion through the joint design of visual and auditory elements, yet most prior work…

EgoSound: Benchmarking Sound Understanding in Egocentric Videos

2026-02-15 · Bingwen Zhu, Yuqian Fu, Qiaole Dong, Guolei Sun 외 arxiv

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating sight, sound, and motion to reason about…

Causal Inference

Interpreting Audiograms with Multi-stage Neural Networks

2021-12-17 · Shufan Li, Congxi Lu, Linkai Li, Jirong Duan 외

Audiograms are a particular type of line charts representing individuals' hearing level at various frequencies. They are used by audiologists to diagnose hearing loss, and further select and tune appropriate hearing aids…

A Classification Model Utilizing Facial Landmark Tracking to Determine Sentence Types for American Sign Language Recognition

2022-11-23 · Janice Nguyen, Y. Curtis Wang

The deaf and hard of hearing community relies on American Sign Language (ASL) as their primary mode of communication, but communication with others who do not know ASL can be difficult, especially during emergencies wher…

Landmark TrackingSentenceSign Language Recognition

Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners

2024-02-27 · CVPR 2024 1 · Yazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang 외

Video and audio content creation serves as the core technique for the movie industry and professional users. Recently, existing diffusion-based methods tackle video and audio generation separately, which hinders the tech…

Audio GenerationDenoising