Transductive Learning with String Kernels for Cross-Domain Text Classification
For many text classification tasks, there is a major problem posed by the lack of labeled data in a target domain. Although classifiers for a target domain can be trained on labeled text data from a related source domain, the accuracy of such classifiers is usually lower in the cross-domain setting. Recently, string kernels have obtained state-of-the-art results in various text classification tasks such as native language identification or automatic essay scoring. Moreover, classifiers based on string kernels have been found to be robust to the distribution gap between different domains. In this paper, we formally describe an algorithm composed of two simple yet effective transductive learning approaches to further improve the results of string kernels in cross-domain settings. By adapting string kernels to the test set without using the ground-truth test labels, we report significantly better accuracy rates in cross-domain English polarity classification.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationCross-Domain Text ClassificationGeneral ClassificationLanguage IdentificationNative Language Identificationtext-classificationText ClassificationTransductive LearningSimilar Papers 제목 키워드 기반
Improving the results of string kernels in sentiment analysis and Arabic dialect identification by adapting them to your test set
Recently, string kernels have obtained state-of-the-art results in various text classification tasks such as Arabic dialect identification or native language identification. In this paper, we apply two simple yet effecti…
Dialect IdentificationGeneral ClassificationLanguage IdentificationNative Language Identification+4Automated essay scoring with string kernels and word embeddings
In this work, we present an approach based on combining string kernels and word embeddings for automatic essay scoring. String kernels capture the similarity among strings based on counting common character n-grams, whic…
Automated Essay ScoringDialect IdentificationGeneral ClassificationLanguage Identification+4Single and Cross-domain Polarity Classification using String Kernels
The polarity classification task aims at automatically identifying whether a subjective text is positive or negative. When the target domain is different from those where a model was trained, we refer to a cross-domain s…
ClassificationDomain AdaptationGeneral ClassificationText ClassificationBOSS: Bayesian Optimization over String Spaces
This article develops a Bayesian optimization (BO) method which acts directly over raw strings, proposing the first uses of string kernels and genetic algorithms within BO loops. Recent applications of BO over strings ha…
Bayesian OptimizationComparison of Short-Text Sentiment Analysis Methods for Croatian
We focus on the task of supervised sentiment classification of short and informal texts in Croatian, using two simple yet effective methods: word embeddings and string kernels. We investigate whether word embeddings offe…
General ClassificationSentiment AnalysisSentiment ClassificationStock Price Prediction+2