Stacked Sentence-Document Classifier Approach for Improving Native Language Identification
In this paper, we describe the approach of the ItaliaNLP Lab team to native language identification and discuss the results we submitted as participants to the essay track of NLI Shared Task 2017. We introduce for the first time a 2-stacked sentence-document architecture for native language identification that is able to exploit both local sentence information and a wide set of general-purpose features qualifying the lexical and grammatical structure of the whole document. When evaluated on the official test set, our sentence-document stacked architecture obtained the best result among all the participants of the essay track with an F1 score of 0.8818.
Code (0)
등록된 구현이 없습니다.
Tasks
Document ClassificationLanguage IdentificationNative Language IdentificationSentenceSentence ClassificationText ClassificationSimilar Papers 제목 키워드 기반
Shortcut-Stacked Sentence Encoders for Multi-Domain Inference
We present a simple sequential sentence encoder for multi-domain natural language inference. Our encoder is based on stacked bidirectional LSTM-RNNs with shortcut connections and fine-tuning of word embeddings. The overa…
Natural Language InferenceSentenceWord EmbeddingsDocument Image Classification with Intra-Domain Transfer Learning and Stacked Generalization of Deep Convolutional Neural Networks
In this work, a region-based Deep Convolutional Neural Network framework is proposed for document structure learning. The contribution of this work involves efficient training of region based classifiers and effective en…
document-image-classificationDocument Image ClassificationGeneral Classificationimage-classification+2An Iterative Approach for Mining Parallel Sentences in a Comparable Corpus
We describe an approach for mining parallel sentences in a collection of documents in two languages. While several approaches have been proposed for doing so, our proposal differs in several respects. First, we use a doc…
Information RetrievalMachine TranslationSentenceTranslationCrosslingual Document Embedding as Reduced-Rank Ridge Regression
There has recently been much interest in extending vector-based word representations to multiple languages, such that words can be compared across languages. In this paper, we shift the focus from words to documents and …
Document EmbeddingregressionRetrievalSentenceRTM Stacking Results for Machine Translation Performance Prediction
We obtain new results using referential translation machines with increased number of learning models in the set of results that are stacked to obtain a better mixture of experts prediction. We combine features extracted…
Machine TranslationMixture-of-ExpertsPredictionSentence+1