paper-with-me

홈 › Papers

Causal Tracing of Audio-Text Fusion in Large Audio Language Models

2026-03-14 · Wei-Chih Chen, Chien-yu Huang, Hung-yi Lee arxiv

Despite the strong performance of large audio language models (LALMs) in various tasks, exactly how and where they integrate acoustic features with textual context remains unclear. We adapt causal tracing to investigate the internal information flow of LALMs during audio comprehension. By conducting layer-wise and token-wise analyses across DeSTA, Qwen, and Voxtral, we evaluate the causal effects of individual hidden states. Layer-wise analysis identifies different fusion strategies, from progressive integration in DeSTA to abrupt late-stage fusion in Qwen. Token-wise analysis shows that the final sequence token acts as an informational bottleneck where the network decisively retrieves relevant information from the audio. We also observe an attention-like query mechanism at intermediate token positions that triggers the model to pull task-relevant audio context. These findings provide a clear characterization of when and where multi-modal integration occurs within LALMs.

📄 PDF Abstract BibTeX arXiv:2603.13768

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion

2025-06-02 · Ajinkya Kulkarni, Sandipana Dowerah, Tanel Alumae, Mathew Magimai. -Doss

Audio deepfakes are acquiring an unprecedented level of realism with advanced AI. While current research focuses on discerning real speech from spoofed speech, tracing the source system is equally crucial. This work prop…

Face SwappingMetric Learning

Localizing and Editing Knowledge in Large Audio-Language Models

2026-03-15 · Sung Kyun Chung, Jiaheng Dong, Qiuchi Hu, Gongping Huang 외 arxiv

Large Audio-Language Models (LALMs) have shown strong performance in speech understanding, making speech a natural interface for accessing factual information. Yet they are trained on static corpora and may encode incorr…

Smark: A Watermark for Text-to-Speech Diffusion Models via Discrete Wavelet Transform

2025-12-21 · Yichuan Zhang, Chengxin Li, Yujie Gu arxiv

Text-to-Speech (TTS) diffusion models generate high-quality speech, which raises challenges for the model intellectual property protection and speech tracing for legal use. Audio watermarking is a promising solution. How…

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming

2026-08-06 · Menglin Han, Yang Ding, Yulei Lu, Haoran Yu 외 arxiv

Real-time long-form avatar audio--video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consistency. Adapting a pretrained bidirectional model to this setting pre…

Video Generation

DiscoForcing: A Unified Framework for Real-Time Audio-Driven Character Control with Diffusion Forcing

2026-05-27 · Kaiyang Ji, Bingsheng Qian, Binghuan Wu, Kangyi Chen 외 arxiv

We study real-time audio-responsive character control as a deployment-faithful problem: strictly causal, bounded-latency streaming that must generate coherent full-body motion at interactive frame rates while the audio c…