paper-with-me

홈 › Papers

Pareto Probing: Trading Off Accuracy for Complexity

2020-10-05 · EMNLP 2020 11 · Tiago Pimentel, Naomi Saphra, Adina Williams, Ryan Cotterell

The question of how to probe contextual word representations for linguistic structure in a way that is both principled and useful has seen significant attention recently in the NLP literature. In our contribution to this discussion, we argue for a probe metric that reflects the fundamental trade-off between probe complexity and performance: the Pareto hypervolume. To measure complexity, we present a number of parametric and non-parametric metrics. Our experiments using Pareto hypervolume as an evaluation metric show that probes often do not conform to our expectations -- e.g., why should the non-contextual fastText representations encode more morpho-syntactic information than the contextual BERT representations? These results suggest that common, simplistic probing tasks, such as part-of-speech labeling and dependency arc labeling, are inadequate to evaluate the linguistic structure encoded in contextual word representations. This leads us to propose full dependency parsing as a probing task. In support of our suggestion that harder probing tasks are necessary, our experiments with dependency parsing reveal a wide gap in syntactic knowledge between contextual and non-contextual representations.

📄 PDF Abstract BibTeX arXiv:2010.02180

Code (1)

rycolab/pareto-probing 공식 구현 pytorch

Tasks

ARCDependency Parsing

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Multi-Head Attention 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Residual Connection 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
fastText fastText embeddings exploit subword information to construct word embeddings. Representations are learnt of character $n$-grams, and words represented as the sum of the…

Similar Papers 제목 키워드 기반

Low-Complexity Probing via Finding Subnetworks

2021-04-08 · NAACL 2021 4 · Steven Cao, Victor Sanh, Alexander M. Rush

The dominant approach in probing neural networks for linguistic properties is to train a new shallow multi-layer perceptron (MLP) on top of the model's internal representations. This approach can detect properties encode…

An Optimal Procedure to Check Pareto-Optimality in House Markets with Single-Peaked Preferences

2020-02-14 · Aurélie Beynier, Nicolas Maudet, Simon Rey, Parham Shams

Recently, the problem of allocating one resource per agent with initial endowments (house markets) has seen a renewed interest: indeed, while in the domain of strict preferences the Top Trading Cycle algorithm is known t…

Multi-Complexity-Loss DNAS for Energy-Efficient and Memory-Constrained Deep Neural Networks

2022-06-01 · Matteo Risso, Alessio Burrello, Luca Benini, Enrico Macii 외

Neural Architecture Search (NAS) is increasingly popular to automatically explore the accuracy versus computational complexity trade-off of Deep Learning (DL) architectures. When targeting tiny edge devices, the main cha…

Neural Architecture Search

The Economic Analysis of the Common Pool Method through the HARA Utility Functions

2024-08-09 · Mu Lin, Di Zhang, Ben Chen, Hang Zheng

Water market is a contemporary marketplace for water trading and is deemed to one of the most efficient instruments to improve the social welfare. In modern water markets, the two widely used trading systems are an impro…

Intra-request branch orchestration for efficient LLM reasoning

2025-09-29 · Weifan Jiang, Rana Shahout, Yilun Du, Michael Mitzenmacher 외 arxiv

Large Language Models (LLMs) increasingly rely on inference-time reasoning algorithms such as chain-of-thought and multi-branch reasoning to improve accuracy on complex tasks. These methods, however, substantially increa…