paper-with-me

홈 › Papers

How to Compute the Probability of a Word

2024-06-20 · Tiago Pimentel, Clara Meister

Language models (LMs) estimate a probability distribution over strings in a natural language; these distributions are crucial for computing perplexity and surprisal in linguistics research. While we are usually concerned with measuring these values for words, most LMs operate over subwords. Despite seemingly straightforward, accurately computing probabilities over one unit given probabilities over the other requires care. Indeed, we show here that many recent linguistic studies have been incorrectly computing these values. This paper derives the correct methods for computing word probabilities, highlighting issues when relying on language models that use beginning-of-word (bow)-marking tokenisers, e.g., the GPT family. Empirically, we show that correcting the widespread bug in probability computations affects measured outcomes in sentence comprehension and lexical optimisation analyses.

📄 PDF Abstract BibTeX arXiv:2406.14561

Code (2)

tpimentelms/probability-of-a-word 공식 구현 pytorch
lacclab/text-metrics pytorch

Tasks

Sentence

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Adam 설명 없음
Linear Warmup With Cosine Annealing Linear Warmup With Cosine Annealing is a learning rate schedule where we increase the learning rate linearly for $n$ updates and then anneal according to a cosine schedule…

Similar Papers 제목 키워드 기반

Secure Bayesian Federated Analytics for Privacy-Preserving Trend Detection

2021-07-28 · Amit Chaulwar, Michael Huth

Federated analytics has many applications in edge computing, its use can lead to better decision making for service provision, product development, and user experience. We propose a Bayesian approach to trend detection i…

Decision MakingEdge-computingPrivacy Preserving

Training Hybrid Language Models by Marginalizing over Segmentations

2019-07-01 · ACL 2019 7 · Edouard Grave, Sainbayar Sukhbaatar, Piotr Bojanowski, Arm Joulin 외

In this paper, we study the problem of hybrid language modeling, that is using models which can predict both characters and larger units such as character ngrams or words. Using such models, multiple potential segmentati…

Language ModelingLanguage Modelling

Quantum Visual Word Sense Disambiguation: Unraveling Ambiguities Through Quantum Inference Model

2025-12-31 · Wenbo Qiao, Peng Zhang, Qinghua Hu arxiv

Visual word sense disambiguation focuses on polysemous words, where candidate images can be easily confused. Traditional methods use classical probability to calculate the likelihood of an image matching each gloss of th…

Word Sense DisambiguationQuantum Machine LearningImage Matching

Online Infix Probability Computation for Probabilistic Finite Automata

2019-07-01 · ACL 2019 7 · Marco Cognetta, Yo-Sub Han, Soon Chan Kwon

Probabilistic finite automata (PFAs) are com- mon statistical language model in natural lan- guage and speech processing. A typical task for PFAs is to compute the probability of all strings that match a query pattern. A…

Language ModelingLanguage Modelling

Watermarking Text Generated by Black-Box Language Models

2023-05-14 · Xi Yang, Kejiang Chen, Weiming Zhang, Chang Liu 외

LLMs now exhibit human-like skills in various fields, leading to worries about misuse. Thus, detecting generated text is crucial. However, passive detection methods are stuck in domain specificity and limited adversarial…

Adversarial RobustnessLanguage ModellingSpecificityText Generation