Approximate kNN Classification for Biomedical Data
We are in the era where the Big Data analytics has changed the way of interpreting the various biomedical phenomena, and as the generated data increase, the need for new machine learning methods to handle this evolution grows. An indicative example is the single-cell RNA-seq (scRNA-seq), an emerging DNA sequencing technology with promising capabilities but significant computational challenges due to the large-scaled generated data. Regarding the classification process for scRNA-seq data, an appropriate method is the k Nearest Neighbor (kNN) classifier since it is usually utilized for large-scale prediction tasks due to its simplicity, minimal parameterization, and model-free nature. However, the ultra-high dimensionality that characterizes scRNA-seq impose a computational bottleneck, while prediction power can be affected by the "Curse of Dimensionality". In this work, we proposed the utilization of approximate nearest neighbor search algorithms for the task of kNN classification in scRNA-seq data focusing on a particular methodology tailored for high dimensional data. We argue that even relaxed approximate solutions will not affect the prediction performance significantly. The experimental results confirm the original assumption by offering the potential for broader applicability.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationGeneral ClassificationPredictionSimilar Papers 제목 키워드 기반
Explaining Black-box Models for Biomedical Text Classification
In this paper, we propose a novel method named Biomedical Confident Itemsets Explanation (BioCIE), aiming at post-hoc explanation of black-box machine learning models for biomedical text classification. Using sources of …
ClassificationGeneral Classificationtext-classificationText ClassificationLearning Entity-Likeness with Multiple Approximate Matches for Biomedical NER
Biomedical Named Entities are complex, so approximate matching has been used to improve entity coverage. However, the usual approximate matching approach fetches only one matching result, which is often noisy. In this wo…
NERLitMC-BERT: transformer-based multi-label classification of biomedical literature with an application on COVID-19 literature curation
The rapid growth of biomedical literature poses a significant challenge for curation and interpretation. This has become more evident during the COVID-19 pandemic. LitCovid, a literature database of COVID-19 related pape…
ArticlesMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONPerformance evaluation of Machine learning algorithms in Biomedical Document Classification
Document classification is a prevalent task in Natural Language Processing (NLP) with a broad range of applications in the biomedical domain. In biomedical engineering categorization of biomedical literature into predefi…
BIG-bench Machine LearningClassificationDocument ClassificationFeature Imitating Networks Enhance The Performance, Reliability And Speed Of Deep Learning On Biomedical Image Processing Tasks
Feature-Imitating-Networks (FINs) are neural networks that are first trained to approximate closed-form statistical features (e.g. Entropy), and then embedded into other networks to enhance their performance. In this wor…
Brain Tumor ClassificationBrain Tumor SegmentationTumor Segmentation