UsingWord Embedding for Cross-Language Plagiarism Detection
This paper proposes to use distributed representation of words (word embeddings) in cross-language textual similarity detection. The main contributions of this paper are the following: (a) we introduce new cross-language similarity detection methods based on distributed representation of words; (b) we combine the different methods proposed to verify their complementarity and finally obtain an overall F1 score of 89.15% for English-French similarity detection at chunk level (88.5% at sentence level) on a very challenging corpus.
Code (0)
등록된 구현이 없습니다.
Tasks
SentenceWord EmbeddingsSimilar Papers 제목 키워드 기반
Enhancing Plagiarism Detection in Marathi with a Weighted Ensemble of TF-IDF and BERT Embeddings for Low-Resource Language Processing
Plagiarism involves using another person's work or concepts without proper attribution, presenting them as original creations. With the growing amount of data communicated in regional languages such as Marathi -- one of …
Paraphrase IdentificationSentenceSentence EmbeddingsEnglish-Arabic Cross-language Plagiarism Detection
The advancement of the web and information technology has contributed to the rapid growth of digital libraries and automatic machine translation tools which easily translate texts from one language into another. These ha…
Machine TranslationSentenceTranslationWord AlignmentDetecting Cross-Lingual Plagiarism Using Simulated Word Embeddings
Cross-lingual plagiarism (CLP) occurs when texts written in one language are translated into a different language and used without acknowledging the original sources. One of the most common methods for detecting CLP requ…
Word EmbeddingsCrossLang: the system of cross-lingual plagiarism detection
Plagiarism and text reuse become more available with the Internet development. Therefore it is important to check scientific papers for the fact of cheating, especially in Academia. Existing systems of plagiarism detecti…
Using Word Embedding for Cross-Language Plagiarism Detection
This paper proposes to use distributed representation of words (word embeddings) in cross-language textual similarity detection. The main contributions of this paper are the following: (a) we introduce new cross-language…
Machine TranslationSentenceWord Embeddings