paper-with-me

홈 › Papers

Thinking While Listening: Simple Test Time Scaling For Audio Classification

2025-09-24 · Prateek Verma, Mert Pilanci arxiv

We propose a framework that enables neural models to "think while listening" to everyday sounds, thereby enhancing audio classification performance. Motivated by recent advances in the reasoning capabilities of large language models, we address two central questions: (i) how can thinking be incorporated into existing audio classification pipelines to enable reasoning in the category space and improve performance, and (ii) can a new architecture be designed from the ground up to support both thinking and test-time scaling? We demonstrate that in both settings, our models exhibit improved classification accuracy. Leveraging test-time scaling, we observe consistent gains as the number of sampled traces increases. Furthermore, we evaluate two open-source reasoning models, GPT-OSS-20B and Qwen3-14B, showing that while such models are capable of zero-shot reasoning, a lightweight approach--retraining only the embedding matrix of a frozen, smaller model like GPT-2--can surpass the performance of billion-parameter text-based reasoning models.

📄 PDF Abstract BibTeX arXiv:2509.19676

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Classification

Similar Papers 제목 키워드 기반

Chronological Thinking in Full-Duplex Spoken Dialogue Language Models

2025-10-02 · Donghang Wu, Haoyang Zhang, Chen Chen, Tianyu Zhang 외 arxiv

Recent advances in spoken dialogue language models (SDLMs) reflect growing interest in shifting from turn-based to full-duplex systems, where the models continuously perceive user speech streams while generating response…

SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models

2025-10-08 · Cheng-Han Chiang, Xiaofei Wang, Linjie Li, Chung-Ching Lin 외 arxiv

Current large language models (LLMs) and spoken language models (SLMs) begin thinking and taking actions only after the user has finished their turn. This prevents the model from interacting during the user's turn and ca…

Does Thinking More always Help? Understanding Test-Time Scaling in Reasoning Models

2025-06-04 · Soumya Suvra Ghosal, Souradip Chakraborty, Avinash Reddy, Yifu Lu 외

Recent trends in test-time scaling for reasoning models (e.g., OpenAI o1, DeepSeek R1) have led to a popular belief that extending thinking traces using prompts like "Wait" or "Let me rethink" can improve performance. Th…

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning

2026-03-18 · Donghang Wu, Tianyu Zhang, Yuxin Li, Hexin Liu 외 arxiv

During conversational interactions, humans subconsciously engage in concurrent thinking while listening to a speaker. Although this internal cognitive processing may not always manifest as explicit linguistic structures,…

Listening Deepfake Detection: A New Perspective Beyond Speaking-Centric Forgery Analysis

2026-04-14 · Miao Liu, Fangda Wei, Jing Wang, Xinyuan Qian arxiv

Existing deepfake detection research has primarily focused on scenarios where the manipulated subject is actively speaking, i.e., generating fabricated content by altering the speaker's appearance or voice. However, in r…

DeepFake Detection