paper-with-me

Papers

Teaching Physical Awareness to LLMs through Sounds

2025-06-10 · Weiguo Wang, Andy Nie, Wenrui Zhou, Yi Kai, Chengchen Hu

Large Language Models (LLMs) have shown remarkable capabilities in text and multimodal processing, yet they fundamentally lack physical awareness--understanding of real-world physical phenomena. In this work, we present ACORN, a framework that teaches LLMs physical awareness through sound, focusing on fundamental physical phenomena like the Doppler effect, multipath effect, and spatial relationships. To overcome data scarcity, ACORN introduce a physics-based simulator combining real-world sound sources with controlled physical channels to generate diverse training data. Using this simulator, we build AQA-PHY, a comprehensive Audio Question-Answer dataset, and propose an audio encoder that processes both magnitude and phase information. By connecting our audio encoder to state-of-the-art LLMs, we demonstrate reasonable results in both simulated and real-world tasks, such as line-of-sight detection, Doppler effect estimation, and Direction-of-Arrival estimation, paving the way for enabling LLMs to understand physical world.

📄 PDF Abstract BibTeX arXiv:2506.08524

Code (0)

등록된 구현이 없습니다.

Tasks

Direction of Arrival Estimation

Similar Papers 제목 키워드 기반

Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples

2025-05-20 · Chun-Yi Kuan, Hung-Yi Lee

Recent advancements in audio-aware large language models (ALLMs) enable them to process and understand audio inputs. However, these models often hallucinate non-existent sound events, reducing their reliability in real-w…

Generative AI in Education: A Study of Educators' Awareness, Sentiments, and Influencing Factors

2024-03-22 · Aashish Ghimire, James Prather, John Edwards

The rapid advancement of artificial intelligence (AI) and the expanding integration of large language models (LLMs) have ignited a debate about their application in education. This study delves into university instructor…

A Sequential Self Teaching Approach for Improving Generalization in Sound Event Recognition

2020-06-30 · ICML 2020 1 · Anurag Kumar, Vamsi Krishna Ithapu

An important problem in machine auditory perception is to recognize and detect sound events. In this paper, we propose a sequential self-teaching approach to learning sounds. Our main proposition is that it is harder to …

Audio ClassificationTransfer Learning

Object-based synthesis of scraping and rolling sounds based on non-linear physical constraints

2021-12-16 · Vinayak Agarwal, Maddie Cusimano, James Traer, Josh Mcdermott

Sustained contact interactions like scraping and rolling produce a wide variety of sounds. Previous studies have explored ways to synthesize these sounds efficiently and intuitively but could not fully mimic the rich str…

TEACHING -- Trustworthy autonomous cyber-physical applications through human-centred intelligence

2021-07-14 · Davide Bacciu, Siranush Akarmazyan, Eric Armengaud, Manlio Bacco 외

This paper discusses the perspective of the H2020 TEACHING project on the next generation of autonomous applications running in a distributed and highly heterogeneous environment comprising both virtual and physical reso…

Federated Learning