A Log-Linear Model for Unsupervised Text Normalization
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModellingLexical NormalizationmodelText NormalizationSimilar Papers 제목 키워드 기반
An Unsupervised Normalization Algorithm for Noisy Text: A Case Study for Information Retrieval and Stance Detection
A large fraction of textual data available today contains various types of 'noise', such as OCR noise in digitized documents, noise due to informal writing style of users on microblogging sites, and so on. To enable task…
Information RetrievalOptical Character Recognition (OCR)RetrievalStance Detection+1Sentiment Analysis in Code-Mixed Telugu-English Text with Unsupervised Data Normalization
In a multilingual society, people communicate in more than one language, leading to Code-Mixed data. Sentimental analysis on Code-Mixed Telugu-English Text (CMTET) poses unique challenges. The unstructured nature of the …
Sentiment AnalysisExperiments in Extractive Summarization: Integer Linear Programming, Term/Sentence Scoring, and Title-driven Models
In this paper, we revisit the challenging problem of unsupervised single-document summarization and study the following aspects: Integer linear programming (ILP) based algorithms, Parameterized normalization of term and …
Document SummarizationExtractive SummarizationSentenceThe Denoised Web Treebank: Evaluating Dependency Parsing under Noisy Input Conditions
We introduce the Denoised Web Treebank: a treebank including a normalization layer and a corresponding evaluation metric for dependency parsing of noisy text, such as Tweets. This benchmark enables the evaluation of pars…
Dependency ParsingLexical NormalizationMachine TranslationText Normalization+1The Edge of Orthogonality: A Simple View of What Makes BYOL Tick
Self-predictive unsupervised learning methods such as BYOL or SimSiam have shown impressive results, and counter-intuitively, do not collapse to trivial representations. In this work, we aim at exploring the simplest pos…