paper-with-me

홈 › Papers

Can persistent homology whiten Transformer-based black-box models? A case study on BERT compression

2023-12-17 · Luis Balderas, Miguel Lastra, José M. Benítez

Large Language Models (LLMs) like BERT have gained significant prominence due to their remarkable performance in various natural language processing tasks. However, they come with substantial computational and memory costs. Additionally, they are essentially black-box models, challenging to explain and interpret. In this article, we propose Optimus BERT Compression and Explainability (OBCE), a methodology to bring explainability to BERT models using persistent homology, aiming to measure the importance of each neuron by studying the topological characteristics of their outputs. As a result, we can compress BERT significantly by reducing the number of parameters (58.47% of the original parameters for BERT Base, 52.3% for BERT Large). We evaluated our methodology on the standard GLUE Benchmark, comparing the results with state-of-the-art techniques and achieving outstanding results. Consequently, our methodology can "whiten" BERT models by providing explainability to its neurons and reducing the model's size, making it more suitable for deployment on resource-constrained devices.

📄 PDF Abstract BibTeX arXiv:2312.10702

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Weight Decay 설명 없음
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Persistent Topology of Syntax

2015-07-18 · Alexander Port, Iulia Gheorghita, Daniel Guth, John M. Clark 외

We study the persistent homology of the data set of syntactic parameters of the world languages. We show that, while homology generators behave erratically over the whole data set, non-trivial persistent homology appears…

Position

A persistent-homology-based turbulence index & some applications of TDA on financial markets

2022-03-10 · Miguel A. Ruiz-Ortiz, José Carlos Gómez-Larrañaga, Jesús Rodríguez-Viorato

Topological Data Analysis (TDA) is a modern approach to Data Analysis focusing on the topological features of data; it has been widely studied in recent years and used extensively in Biology, Physics, and many other area…

Time SeriesTopological Data Analysis

Persistent Intersection Homology for the Analysis of Discrete Data

2019-07-31 · Bastian Rieck, Markus Banagl, Filip Sadlo, Heike Leitte

Topological data analysis is becoming increasingly relevant to support the analysis of unstructured data sets. A common assumption in data analysis is that the data set is a sample---not necessarily a uniform one---of so…

Topological Data Analysis

Persistent Homology for High-dimensional Data Based on Spectral Methods

2023-11-06 · Sebastian Damrich, Philipp Berens, Dmitry Kobak

Persistent homology is a popular computational tool for analyzing the topology of point clouds, such as the presence of loops or voids. However, many real-world datasets with low intrinsic dimensionality reside in an amb…

Persistent Homology and Graphs Representation Learning

2021-02-25 · Mustafa Hajij, Ghada Zamzmi, Xuanting Cai

This article aims to study the topological invariant properties encoded in node graph representational embeddings by utilizing tools available in persistent homology. Specifically, given a node embedding representation a…

Representation Learning