paper-with-me

홈 › Papers

Lightweight Baselines for Medical Abstract Classification: DistilBERT with Cross-Entropy as a Strong Default

2025-10-11 · Jiaqi Liu, Tong Wang, Su Liu, Xin Hu, Ran Tong, Lanruo Wang, Jiexi Xu arxiv

The research evaluates lightweight medical abstract classification methods to establish their maximum performance capabilities under financial budget restrictions. On the public medical abstracts corpus, we finetune BERT base and Distil BERT with three objectives cross entropy (CE), class weighted CE, and focal loss under identical tokenization, sequence length, optimizer, and schedule. DistilBERT with plain CE gives the strongest raw argmax trade off, while a post hoc operating point selection (validation calibrated, classwise thresholds) sub stantially improves deployed performance; under this tuned regime, focal benefits most. We report Accuracy, Macro F1, and WeightedF1, release evaluation artifacts, and include confusion analyses to clarify error structure. The practical takeaway is to start with a compact encoder and CE, then add lightweight calibration or thresholding when deployment requires higher macro balance.

📄 PDF Abstract BibTeX arXiv:2510.10025

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Automated Text Mining of Experimental Methodologies from Biomedical Literature

2024-04-21 · Ziqing Guo

Biomedical literature is a rapidly expanding field of science and technology. Classification of biomedical texts is an essential part of biomedicine research, especially in the field of biology. This work proposes the fi…

ArticlesClassificationLanguage ModelingLanguage Modelling

Assessment of DistilBERT performance on Named Entity Recognition task for the detection of Protected Health Information and medical concepts

2020-11-01 · EMNLP (ClinicalNLP) 2020 11 · Macarious Abadeer

Bidirectional Encoder Representations from Transformers (BERT) models achieve state-of-the-art performance on a number of Natural Language Processing tasks. However, their model size on disk often exceeds 1 GB and the pr…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)

Analyzing the Generalizability of Deep Contextualized Language Representations For Text Classification

2023-03-22 · Berfu Buyukoz

This study evaluates the robustness of two state-of-the-art deep contextual language representations, ELMo and DistilBERT, on supervised learning of binary protest news classification and sentiment analysis of product re…

Binary text classificationNews ClassificationSentiment Analysistext-classification+1

Comparative Efficiency Analysis of Lightweight Transformer Models: A Multi-Domain Empirical Benchmark for Enterprise NLP Deployment

2026-01-01 · Muhammad Shahmeer Khan arxiv

In the rapidly evolving landscape of enterprise natural language processing (NLP), the demand for efficient, lightweight models capable of handling multi-domain text automation tasks has intensified. This study conducts …

Hyperparameter OptimizationHate Speech Detection

Sentiment Analysis Of Shopee Product Reviews Using Distilbert

2025-11-27 · Zahri Aksa Dautd, Aviv Yuniar Rahman arxiv

The rapid growth of digital commerce has led to the accumulation of a massive number of consumer reviews on online platforms. Shopee, as one of the largest e-commerce platforms in Southeast Asia, receives millions of pro…

Sentiment Analysis