Using Word Embedding for Cross-Language Plagiarism Detection
This paper proposes to use distributed representation of words (word embeddings) in cross-language textual similarity detection. The main contributions of this paper are the following: (a) we introduce new cross-language similarity detection methods based on distributed representation of words; (b) we combine the different methods proposed to verify their complementarity and finally obtain an overall F1 score of 89.15{\%} for English-French similarity detection at chunk level (88.5{\%} at sentence level) on a very challenging corpus.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationSentenceWord EmbeddingsSimilar Papers 제목 키워드 기반
English-Arabic Cross-language Plagiarism Detection
The advancement of the web and information technology has contributed to the rapid growth of digital libraries and automatic machine translation tools which easily translate texts from one language into another. These ha…
Machine TranslationSentenceTranslationWord AlignmentUsingWord Embedding for Cross-Language Plagiarism Detection
This paper proposes to use distributed representation of words (word embeddings) in cross-language textual similarity detection. The main contributions of this paper are the following: (a) we introduce new cross-language…
SentenceWord EmbeddingsDetecting Cross-Lingual Plagiarism Using Simulated Word Embeddings
Cross-lingual plagiarism (CLP) occurs when texts written in one language are translated into a different language and used without acknowledging the original sources. One of the most common methods for detecting CLP requ…
Word EmbeddingsA Resource-Light Method for Cross-Lingual Semantic Textual Similarity
Recognizing semantically similar sentences or paragraphs across languages is beneficial for many tasks, ranging from cross-lingual information retrieval and plagiarism detection to machine translation. Recently proposed …
Cross-Lingual Information RetrievalCross-Lingual Semantic Textual SimilarityInformation RetrievalMachine Translation+9Enhancing Plagiarism Detection in Marathi with a Weighted Ensemble of TF-IDF and BERT Embeddings for Low-Resource Language Processing
Plagiarism involves using another person's work or concepts without proper attribution, presenting them as original creations. With the growing amount of data communicated in regional languages such as Marathi -- one of …
Paraphrase IdentificationSentenceSentence Embeddings