paper-with-me

홈 › Papers

Inside the Latent Flow: Causal Deciphering of Attention Dynamics in Audio Separation Foundation Models

2026-06-08 · Yuxuan Chen, Haoyuan Yu, Peize He arxiv

Flow-matching transformers achieve strong audio separation, yet their attention dynamics are opaque. We adapt established causal-intervention principles into a deterministic, inference-time probing protocol for SAM Audio. Orthogonal probing uncovers a dual-pathway text-conditioning mechanism: additive injections control semantic identity, while cross-attention refines acoustic structure. We observe an asynchronous layerwise convergence: stable layers build temporal scaffolds early, whereas fast layers continue resolving artifacts during sampling. The model also attenuates temporal segmentation cues to maintain continuous-flow stability. Using these insights, we propose Layer-Selective Attention Caching (LSAC), a training-free acceleration method that caches attention in stable layers. Across acoustic complexities, LSAC cuts self-attention computation by about ~25% with negligible quality loss and yields up to 6.7x higher quality retention than naive step reduction.

📄 PDF Abstract BibTeX arXiv:2606.10046

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

InfoFlowNet: A Multi-head Attention-based Self-supervised Learning Model with Surrogate Approach for Uncovering Brain Effective Connectivity

2023-11-30 · Chun-Hsiang Chuang, Shao-Xun Fang, Chih-Sheng Huang, Weiping Ding

Deciphering brain network topology can enhance the depth of neuroscientific knowledge and facilitate the development of neural engineering methods. Effective connectivity, a measure of brain network dynamics, is particul…

Causal DiscoveryCausal InferenceEEGElectroencephalogram (EEG)+2

Deciphering interventional dynamical causality from non-intervention systems

2024-06-29 · Jifan Shi, Yang Li, Juan Zhao, Siyang Leng 외

Detecting and quantifying causality is a focal topic in the fields of science, engineering, and interdisciplinary studies. However, causal studies on non-intervention systems attract much attention but remain extremely c…

Time Series

Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality

2024-10-07 · Guanyu Zhou, Yibo Yan, Xin Zou, Kun Wang 외

Multimodal Large Language Models (MLLMs) have emerged as a central focus in both industry and academia, but often suffer from biases introduced by visual and language priors, which can lead to multimodal hallucination. T…

Causal InferencecounterfactualCounterfactual ReasoningHallucination+4

Latent Reasoning with Normalizing Flows

2026-06-04 · Guancheng Tu, Xiangjun Fu, Suhao Yu, Yao Tang 외 arxiv

Large language models often improve reasoning by generating explicit chain-of-thought (CoT), demonstrating the importance of intermediate computation. However, textual CoT forces this computation through a discrete, seri…

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

2026-06-11 · Guozhen Zhang, Xuerui Qiu, Yutao Cui, Tianhui Song 외 arxiv

Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space. In this paper, we present HYDRA-X, the first UMM that unifies image an…