Multi-View Unsupervised User Feature Embedding for Social Media-based Substance Use Prediction
In this paper, we demonstrate how the state-of-the-art machine learning and text mining techniques can be used to build effective social media-based substance use detection systems. Since a substance use ground truth is difficult to obtain on a large scale, to maximize system performance, we explore different unsupervised feature learning methods to take advantage of a large amount of unsupervised social media data. We also demonstrate the benefit of using multi-view unsupervised feature learning to combine heterogeneous user information such as Facebook {`}likes{''} and {`}status updates{''} to enhance system performance. Based on our evaluation, our best models achieved 86{\%} AUC for predicting tobacco use, 81{\%} for alcohol use and 84{\%} for illicit drug use, all of which significantly outperformed existing methods. Our investigation has also uncovered interesting relations between a user{'}s social media behavior (e.g., word usage) and substance use.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Exploring Instance Relations for Unsupervised Feature Embedding
Despite the great progress achieved in unsupervised feature embedding, existing contrastive learning methods typically pursue view-invariant representations through attracting positive sample pairs and repelling negative…
Contrastive Learningimage-classificationImage ClassificationRelation+1Adaptive Similarity Embedding for Unsupervised Multi-View Feature Selection
Multi-view learning has become a significant research topic in image processing, data mining, and machine learning due to the proliferation of multi-view data. Considering the difficulty in obtaining labeled data in man…
feature selectionMULTI-VIEW LEARNINGExploring the Value of Multi-View Learning for Session-Aware Query Representation
Recent years have witnessed a growing interest towards learning distributed query representations that are able to capture search intent semantics. Most existing approaches learn query embeddings using relevance supervis…
Document RankingMULTI-VIEW LEARNINGRepresentation LearningRetrievalExploring the Value of Multi-View Learning for Session-Aware Query Representation
Recent years have witnessed a growing interest towards learning distributed query representations that are able to capture search intent semantics. Most existing approaches learn query embeddings using relevance supervis…
Document RankingMULTI-VIEW LEARNINGRepresentation LearningRetrievalExploring Word Embeddings for Unsupervised Textual User-Generated Content Normalization
Text normalization techniques based on rules, lexicons or supervised training requiring large corpora are not scalable nor domain interchangeable, and this makes them unsuitable for normalizing user-generated content (UG…
Semantic SimilaritySemantic Textual SimilarityText NormalizationWord Embeddings