paper-with-me

Papers

Mediators in Determining what Processing BERT Performs First

2021-04-13 · NAACL 2021 4 · Aviv Slobodkin, Leshem Choshen, Omri Abend

Probing neural models for the ability to perform downstream tasks using their activation patterns is often used to localize what parts of the network specialize in performing what tasks. However, little work addressed potential mediating factors in such comparisons. As a test-case mediating factor, we consider the prediction's context length, namely the length of the span whose processing is minimally required to perform the prediction. We show that not controlling for context length may lead to contradictory conclusions as to the localization patterns of the network, depending on the distribution of the probing dataset. Indeed, when probing BERT with seven tasks, we find that it is possible to get 196 different rankings between them when manipulating the distribution of context lengths in the probing dataset. We conclude by presenting best practices for conducting such comparisons in the future.

📄 PDF Abstract BibTeX arXiv:2104.06400

Code (1)

lovodkin93/BERT-context-distance 공식 구현

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Weight Decay 설명 없음
Adam 설명 없음

Similar Papers 제목 키워드 기반

Mediators: Conversational Agents Explaining NLP Model Behavior

2022-06-13 · Nils Feldhus, Ajay Madhavan Ravichandran, Sebastian Möller

The human-centric explainable artificial intelligence (HCXAI) community has raised the need for framing the explanation process as a conversation between human and machine. In this position paper, we establish desiderata…

Explainable artificial intelligencemodelPositionSentiment Analysis

SMaRT: Online Reusable Resource Assignment and an Application to Mediation in the Kenyan Judiciary

2026-02-20 · Shafkat Farabi, Didac Marti Pinto, Wei Lu, Manuel Ramos-Maqueda 외 arxiv

Motivated by the problem of assigning mediators to cases in the Kenyan judicial system, we study an online resource allocation problem where incoming tasks (cases) must be immediately assigned to available, capacity-cons…

Incorporating Causal Analysis into Diversified and Logical Response Generation

2022-09-20 · Jiayi Liu, Wei Wei, Zhixuan Chu, Xing Gao 외

Although the Conditional Variational AutoEncoder (CVAE) model can generate more diversified responses than the traditional Seq2Seq model, the responses often have low relevance with the input words or are illogical with …

Response Generation

Incorporating Casual Analysis into Diversified and Logical Response Generation

2022-10-01 · COLING 2022 10 · Jiayi Liu, Wei Wei, Zhixuan Chu, Xing Gao 외

Although the Conditional Variational Auto-Encoder (CVAE) model can generate more diversified responses than the traditional Seq2Seq model, the responses often have low relevance with the input words or are illogical with…

Response Generation

Bayesian Persuasion with Mediators

2022-03-08 · Itai Arieli, Yakov Babichenko, Fedor Sandomirskiy

An informed sender communicates with an uninformed receiver through a sequence of uninformed mediators; agents' utilities depend on receiver's action and the state. For any number of mediators, the sender's optimal value…