Utilizing Semantic Textual Similarity for Clinical Survey Data Feature Selection
Survey data can contain a high number of features while having a comparatively low quantity of examples. Machine learning models that attempt to predict outcomes from survey data under these conditions can overfit and result in poor generalizability. One remedy to this issue is feature selection, which attempts to select an optimal subset of features to learn upon. A relatively unexplored source of information in the feature selection process is the usage of textual names of features, which may be semantically indicative of which features are relevant to a target outcome. The relationships between feature names and target names can be evaluated using language models (LMs) to produce semantic textual similarity (STS) scores, which can then be used to select features. We examine the performance using STS to select features directly and in the minimal-redundancy-maximal-relevance (mRMR) algorithm. The performance of STS as a feature selection metric is evaluated against preliminary survey data collected as a part of a clinical study on persistent post-surgical pain (PPSP). The results suggest that features selected with STS can result in higher performance models compared to traditional feature selection algorithms.
Code (1)
Tasks
feature selectionSemantic Textual SimilaritySTSSurveyMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
MedSTS: A Resource for Clinical Semantic Textual Similarity
The wide adoption of electronic health records (EHRs) has enabled a wide range of applications leveraging EHR data. However, the meaningful use of EHR data largely depends on our ability to efficiently extract and consol…
Decision MakingSemantic SimilaritySemantic Textual SimilaritySentence+1Evaluating the Utility of Model Configurations and Data Augmentation on Clinical Semantic Textual Similarity
In this paper, we apply pre-trained language models to the Semantic Textual Similarity (STS) task, with a specific focus on the clinical domain. In low-resource setting of clinical STS, these large models tend to be impr…
Data AugmentationSemantic Textual SimilaritySTSAdvances and Challenges in Semantic Textual Similarity: A Comprehensive Survey
Semantic Textual Similarity (STS) research has expanded rapidly since 2021, driven by advances in transformer architectures, contrastive learning, and domain-specific techniques. This survey reviews progress across six k…
Semantic Textual SimilarityContrastive LearningCLISTER : A Corpus for Semantic Textual Similarity in French Clinical Narratives
Modern Natural Language Processing relies on the availability of annotated corpora for training and evaluating models. Such resources are scarce, especially for specialized domains in languages other than English. In par…
Semantic SimilaritySemantic Textual SimilaritySentenceSentence Embeddings+3Linking Symptom Inventories using Semantic Textual Similarity
An extensive library of symptom inventories has been developed over time to measure clinical symptoms, but this variety has led to several long standing issues. Most notably, results drawn from different settings and stu…
Decision MakingSemantic Textual SimilaritySTS