paper-with-me

홈 › Papers

An Information Extraction Study: Take In Mind the Tokenization!

2023-03-27 · Christos Theodoropoulos, Marie-Francine Moens

Current research on the advantages and trade-offs of using characters, instead of tokenized text, as input for deep learning models, has evolved substantially. New token-free models remove the traditional tokenization step; however, their efficiency remains unclear. Moreover, the effect of tokenization is relatively unexplored in sequence tagging tasks. To this end, we investigate the impact of tokenization when extracting information from documents and present a comparative study and analysis of subword-based and character-based models. Specifically, we study Information Extraction (IE) from biomedical texts. The main outcome is twofold: tokenization patterns can introduce inductive bias that results in state-of-the-art performance, and the character-based models produce promising results; thus, transitioning to token-free IE models is feasible.

📄 PDF Abstract BibTeX arXiv:2303.15100

Code (1)

christos42/inductive_bias_IE 공식 구현 pytorch

Tasks

Inductive BiasNamed Entity Recognition (NER)Relation Extraction

Similar Papers 제목 키워드 기반

A Study of Feature Extraction techniques for Sentiment Analysis

2019-06-04 · Avinash Madasu, Sivasankar E

Sentiment Analysis refers to the study of systematically extracting the meaning of subjective text . When analysing sentiments from the subjective text using Machine Learning techniques,feature extraction becomes a signi…

Sentiment Analysis

Spectral Gaps and Spatial Priors: Studying Hyperspectral Downstream Adaptation Using TerraMind

2026-03-04 · Julia Anna Leonardi, Johannes Jakubik, Paolo Fraccaro, Maria Antonia Brovelli arxiv

Geospatial Foundation Models (GFMs) typically lack native support for Hyperspectral Imaging (HSI) due to the complexity and sheer size of high-dimensional spectral data. This study investigates the adaptability of TerraM…

Mind the Gap: A Closer Look at Tokenization for Multiple-Choice Question Answering with LLMs

2025-09-18 · Mario Sanz-Guerrero, Minh Duc Bui, Katharina von der Wense arxiv

When evaluating large language models (LLMs) with multiple-choice question answering (MCQA), it is common to end the prompt with the string "Answer:" to facilitate automated answer extraction via next-token probabilities…

Question Answering

memorAIs: an Optical Character Recognition and Rule-Based Medication Intake Reminder-Generating Solution

2023-12-11 · Eden Shaveet, Utkarsh Singh, Nicholas Assaderaghi, Maximo Librandi

Memory-based medication non-adherence is an unsolved problem that is responsible for considerable disease burden in the United States. Digital medication intake reminder solutions with minimal onboarding requirements tha…

FrictionOptical Character Recognition

DUFormer: Solving Power Line Detection Task in Aerial Images using Semantic Segmentation

2023-04-12 · Deyu An, Qiang Zhang, Jianshu Chao, Ting Li 외

Unmanned aerial vehicles (UAVs) are frequently used for inspecting power lines and capturing high-resolution aerial images. However, detecting power lines in aerial images is difficult,as the foreground data(i.e, power l…

Inductive BiasLine DetectionSegmentationSemantic Segmentation