Heavy-tailed Representations, Text Polarity Classification & Data Augmentation
The dominant approaches to text representation in natural language rely on learning embeddings on massive corpora which have convenient properties such as compositionality and distance preservation. In this paper, we develop a novel method to learn a heavy-tailed embedding with desirable regularity properties regarding the distributional tails, which allows to analyze the points far away from the distribution bulk using the framework of multivariate extreme value theory. In particular, a classifier dedicated to the tails of the proposed embedding is obtained which performance outperforms the baseline. This classifier exhibits a scale invariance property which we leverage by introducing a novel text generation method for label preserving dataset augmentation. Numerical experiments on synthetic and real text data demonstrate the relevance of the proposed framework and confirm that this method generates meaningful sentences with controllable attribute, e.g. positive or negative sentiment.
Code (0)
등록된 구현이 없습니다.
Tasks
AttributeClassificationData AugmentationGeneral ClassificationSentiment AnalysisText ClassificationText GenerationSimilar Papers 제목 키워드 기반
Single-Stage Heavy-Tailed Food Classification
Deep learning based food image classification has enabled more accurate nutrition content analysis for image-based dietary assessment by predicting the types of food in eating occasion images. However, there are two majo…
Classificationimage-classificationImage ClassificationNutritionERNIE-NLI: Analyzing the Impact of Domain-Specific External Knowledge on Enhanced Representations for NLI
We examine the effect of domain-specific external knowledge variations on deep large scale language model performance. Recent work in enhancing BERT with external knowledge has been very popular, resulting in models such…
Language ModelingLanguage ModellingNatural Language InferenceHeavy-Tailed Process Priors for Selective Shrinkage
Heavy-tailed distributions are often used to enhance the robustness of regression and classification methods to outliers in output space. Often, however, we are confronted with ``outliers'' in input space, which are iso…
Gaussian ProcessesGeneral ClassificationregressionA Multi-View Sentiment Corpus
Sentiment Analysis is a broad task that involves the analysis of various aspect of the natural language text. However, most of the approaches in the state of the art usually investigate independently each aspect, i.e. Su…
Emotion RecognitionGeneral ClassificationOpinion MiningSentiment AnalysisAsymptotic Classification Error for Heavy-Tailed Renewal Processes
Despite the widespread occurrence of classification problems and the increasing collection of point process data across many disciplines, study of error probability for point process classification only emerged very rece…
Classification