paper-with-me

홈 › Papers

BioUNER: A Benchmark Dataset for Clinical Urdu Named Entity Recognition

2026-04-03 · Wazir Ali, Adeeb Noor, Sanaullah Mahar, Alia, Muhammad Mazhar Younas arxiv

In this article, we present a gold-standard benchmark dataset for Biomedical Urdu Named Entity Recognition (BioUNER), developed by crawling health-related articles from online Urdu news portals, medical prescriptions, and hospital health blogs and websites. After preprocessing, three native annotators with familiarity in the medical domain participated in the annotation process using the Doccano text annotation tool and annotated 153K tokens. Following annotation, the proposed BioiUNER dataset was evaluated both intrinsically and extrinsically. An inter-annotator agreement score of 0.78 was achieved, thereby validating the dataset as gold-standard quality. To demonstrate the utility and benchmarking capability of the dataset, we evaluated several machine learning and deep learning models, including Support Vector Machines (SVM), Long Short-Term Memory networks (LSTM), Multilingual BERT (mBERT), and XLM-RoBERTa. The gold-standard BioUNER dataset serves as a reliable benchmark and a valuable addition to Urdu language processing resources.

📄 PDF Abstract BibTeX arXiv:2604.02904

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Benchmark Dataset and a Framework for Urdu Multimodal Named Entity Recognition

2025-05-08 · Hussain Ahmad, Qingyang Zeng, Jing Wan

The emergence of multimodal content, particularly text and images on social media, has positioned Multimodal Named Entity Recognition (MNER) as an increasingly important area of research within Natural Language Processin…

named-entity-recognitionNamed Entity Recognition

Benchmark Performance of Machine And Deep Learning Based Methodologies for Urdu Text Document Classification

2020-03-03 · Muhammad Nabeel Asim, Muhammad Usman Ghani, Muhammad Ali Ibrahim, Sheraz Ahmad 외

In order to provide benchmark performance for Urdu text document classification, the contribution of this paper is manifold. First, it pro-vides a publicly available benchmark dataset manually tagged against 6 classes. S…

Automated Feature EngineeringBIG-bench Machine LearningClassificationDeep Learning+7

Multilingual Hematology Visual Question Answering Dataset

2026-06-24 · Hajra Malik, Hafiza Tooba Aftab, Abdul Rehman, Mohsen Ali 외 arxiv

Vision Language Models (VLMs) have shown promising capabilities in medical image analysis by jointly understanding visual and textual information for tasks such as Visual Question Answering. However, existing hematology …

Visual Question Answering

Hierarchical Text Classification of Urdu News using Deep Neural Network

2021-07-07 · Taimoor Ahmed Javed, Waseem Shahzad, Umair Arshad

Digital text is increasing day by day on the internet. It is very challenging to classify a large and heterogeneous collection of data, which require improved information processing methods to organize text. To classify …

Classificationtext-classificationText Classification

Named Entity Recognition System for Urdu

2012-12-01 · COLING 2012 12 · UmrinderPal Singh, Vishal Goyal, Gurpreet Singh Lehal
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Question Answering