paper-with-me

Papers

Comparing Variation in Tokenizer Outputs Using a Series of Problematic and Challenging Biomedical Sentences

2023-05-15 · Christopher Meaney, Therese A Stukel, Peter C Austin, Michael Escobar

Background & Objective: Biomedical text data are increasingly available for research. Tokenization is an initial step in many biomedical text mining pipelines. Tokenization is the process of parsing an input biomedical sentence (represented as a digital character sequence) into a discrete set of word/token symbols, which convey focused semantic/syntactic meaning. The objective of this study is to explore variation in tokenizer outputs when applied across a series of challenging biomedical sentences. Method: Diaz [2015] introduce 24 challenging example biomedical sentences for comparing tokenizer performance. In this study, we descriptively explore variation in outputs of eight tokenizers applied to each example biomedical sentence. The tokenizers compared in this study are the NLTK white space tokenizer, the NLTK Penn Tree Bank tokenizer, Spacy and SciSpacy tokenizers, Stanza/Stanza-Craft tokenizers, the UDPipe tokenizer, and R-tokenizers. Results: For many examples, tokenizers performed similarly effectively; however, for certain examples, there were meaningful variation in returned outputs. The white space tokenizer often performed differently than other tokenizers. We observed performance similarities for tokenizers implementing rule-based systems (e.g. pattern matching and regular expressions) and tokenizers implementing neural architectures for token classification. Oftentimes, the challenging tokens resulting in the greatest variation in outputs, are those words which convey substantive and focused biomedical/clinical meaning (e.g. x-ray, IL-10, TCR/CD3, CD4+ CD8+, and (Ca2+)-regulated). Conclusion: When state-of-the-art, open-source tokenizers from Python and R were applied to a series of challenging biomedical example sentences, we observed subtle variation in the returned outputs.

📄 PDF Abstract BibTeX arXiv:2305.08787

Code (0)

등록된 구현이 없습니다.

Tasks

Sentencetoken-classificationToken Classification

Similar Papers 제목 키워드 기반

How Long Is a Piece of String? A Brief Empirical Analysis of Tokenizers

2026-01-16 · Jonathan Roberts, Kai Han, Samuel Albanie arxiv

Frontier LLMs are increasingly utilised across academia, society and industry. A commonly used unit for comparing models, their inputs and outputs, and estimating inference pricing is the token. In general, tokens are us…

ModelRadar: Aspect-based Forecast Evaluation

2025-03-31 · Vitor Cerqueira, Luis Roque, Carlos Soares

Accurate evaluation of forecasting models is essential for ensuring reliable predictions. Current practices for evaluating and comparing forecasting models focus on summarising performance into a single score, using metr…

Time SeriesTime Series ForecastingUnivariate Time Series Forecasting

Analyzing Customer-Facing Vendor Experiences with Time Series Forecasting and Monte Carlo Techniques

2024-07-30 · Vivek Kaushik, Jason Tang

eBay partners with external vendors, which allows customers to freely select a vendor to complete their eBay experiences. However, vendor outages can hinder customer experiences. Consequently, eBay can disable a problema…

Time SeriesTime Series Forecasting

Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models

2024-05-08 · Sander Land, Max Bartolo

The disconnect between tokenizer creation and model training in language models allows for specific inputs, such as the infamous SolidGoldMagikarp token, to induce unwanted model behaviour. Although such `glitch tokens',…

Language ModelingLanguage ModellingLarge Language Model

LiPCoT: Linear Predictive Coding based Tokenizer for Self-supervised Learning of Time Series Data via Language Models

2024-08-14 · Md Fahim Anjum

Language models have achieved remarkable success in various natural language processing tasks. However, their application to time series data, a crucial component in many domains, remains limited. This paper proposes LiP…

EEGSelf-Supervised LearningTime Series