paper-with-me

홈 › Papers

Using Word Embedding for Cross-Language Plagiarism Detection

2017-04-01 · EACL 2017 4 · J{\'e}r{\'e}my Ferrero, Laurent Besacier, Didier Schwab, Fr{\'e}d{\'e}ric Agn{\`e}s

This paper proposes to use distributed representation of words (word embeddings) in cross-language textual similarity detection. The main contributions of this paper are the following: (a) we introduce new cross-language similarity detection methods based on distributed representation of words; (b) we combine the different methods proposed to verify their complementarity and finally obtain an overall F1 score of 89.15{\%} for English-French similarity detection at chunk level (88.5{\%} at sentence level) on a very challenging corpus.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentenceWord Embeddings

Similar Papers 제목 키워드 기반

English-Arabic Cross-language Plagiarism Detection

2021-09-01 · RANLP 2021 9 · Naif Alotaibi, Mike Joy

The advancement of the web and information technology has contributed to the rapid growth of digital libraries and automatic machine translation tools which easily translate texts from one language into another. These ha…

Machine TranslationSentenceTranslationWord Alignment

UsingWord Embedding for Cross-Language Plagiarism Detection

2017-02-10 · J. Ferrero, F. Agnes, L. Besacier, D. Schwab

This paper proposes to use distributed representation of words (word embeddings) in cross-language textual similarity detection. The main contributions of this paper are the following: (a) we introduce new cross-language…

SentenceWord Embeddings

Detecting Cross-Lingual Plagiarism Using Simulated Word Embeddings

2017-12-29 · Victor Thompson

Cross-lingual plagiarism (CLP) occurs when texts written in one language are translated into a different language and used without acknowledging the original sources. One of the most common methods for detecting CLP requ…

Word Embeddings

A Resource-Light Method for Cross-Lingual Semantic Textual Similarity

2018-01-19 · Goran Glavaš, Marc Franco-Salvador, Simone Paolo Ponzetto, Paolo Rosso

Recognizing semantically similar sentences or paragraphs across languages is beneficial for many tasks, ranging from cross-lingual information retrieval and plagiarism detection to machine translation. Recently proposed …

Cross-Lingual Information RetrievalCross-Lingual Semantic Textual SimilarityInformation RetrievalMachine Translation+9

Enhancing Plagiarism Detection in Marathi with a Weighted Ensemble of TF-IDF and BERT Embeddings for Low-Resource Language Processing

2025-01-09 · Atharva Mutsaddi, Aditya Choudhary

Plagiarism involves using another person's work or concepts without proper attribution, presenting them as original creations. With the growing amount of data communicated in regional languages such as Marathi -- one of …

Paraphrase IdentificationSentenceSentence Embeddings