paper-with-me

홈 › Papers

LaTIM: Measuring Latent Token-to-Token Interactions in Mamba Models

2025-02-21 · Hugo Pitorro, Marcos Treviso

State space models (SSMs), such as Mamba, have emerged as an efficient alternative to transformers for long-context sequence modeling. However, despite their growing adoption, SSMs lack the interpretability tools that have been crucial for understanding and improving attention-based architectures. While recent efforts provide insights into Mamba's internal mechanisms, they do not explicitly decompose token-wise contributions, leaving gaps in understanding how Mamba selectively processes sequences across layers. In this work, we introduce LaTIM, a novel token-level decomposition method for both Mamba-1 and Mamba-2 that enables fine-grained interpretability. We extensively evaluate our method across diverse tasks, including machine translation, copying, and retrieval-based generation, demonstrating its effectiveness in revealing Mamba's token-to-token interaction patterns.

📄 PDF Abstract BibTeX arXiv:2502.15612

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationMambaRetrievalState Space Models

Methods 이 논문이 사용한 방법론

Mamba Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module.…

Similar Papers 제목 키워드 기반

TabulaTime: A Novel Multimodal Deep Learning Framework for Advancing Acute Coronary Syndrome Prediction through Environmental and Clinical Data Integration

2025-02-24 · Xin Zhang, Liangxiu Han, Stephen White, Saad Hassan 외

Acute Coronary Syndromes (ACS), including ST-segment elevation myocardial infarctions (STEMI) and non-ST-segment elevation myocardial infarctions (NSTEMI), remain a leading cause of mortality worldwide. Traditional cardi…

Data IntegrationFeature EngineeringFeature ImportanceMultimodal Deep Learning+1

Measuring the Mixing of Contextual Information in the Transformer

2022-03-08 · Javier Ferrando, Gerard I. Gállego, Marta R. Costa-jussà

The Transformer architecture aggregates input information through the self-attention mechanism, but there is no clear understanding of how this information is mixed across the entire model. Additionally, recent works hav…

Bias Neutralization Framework: Measuring Fairness in Large Language Models with Bias Intelligence Quotient (BiQ)

2024-04-28 · Malur Narayan, John Pasmore, Elton Sampaio, Vijay Raghavan 외

The burgeoning influence of Large Language Models (LLMs) in shaping public discourse and decision-making underscores the imperative to address inherent biases within these AI systems. In the wake of AI's expansive integr…

Decision MakingFairnessLanguage ModelingLanguage Modelling+1

More Than Words: Collocation Tokenization for Latent Dirichlet Allocation Models

2021-08-24 · Jin Cheevaprawatdomrong, Alexandra Schofield, Attapol T. Rutherford

Traditionally, Latent Dirichlet Allocation (LDA) ingests words in a collection of documents to discover their latent topics using word-document co-occurrences. However, it is unclear how to achieve the best results for l…

Generic Triple-Latent Compression with Gated Associative Retrieval

2026-04-17 · Liu Xiao arxiv

We study generic triple-latent sequence models that maintain a running token state and compressed pair-memory pathway to capture higher-order token interactions without benchmark-specific parsing. The triple-latent famil…