zNLP: Identifying Parallel Sentences in Chinese-English Comparable Corpora
This paper describes the zNLP system for the BUCC 2017 shared task. Our system identifies parallel sentence pairs in Chinese-English comparable corpora by translating word-by-word Chinese sentences into English, using the search engine Solr to select near-parallel sentences and then by using an SVM classifier to identify true parallel sentences from the previous results. It obtains an F1-score of 45{\%} (resp. 32{\%}) on the test (training) set.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationSentenceMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Identify Bilingual Patterns and Phrases from a Bilingual Sentence Pair
This paper presents a method for automatically identifying bilingual grammar patterns and extracting bilingual phrase instances from a given English-Chinese sentence pair. In our approach, the English-Chinese sentence pa…
Machine TranslationSentenceTranslationUM-Corpus: A Large English-Chinese Parallel Corpus for Statistical Machine Translation
Parallel corpus is a valuable resource for cross-language information retrieval and data-driven natural language processing systems, especially for Statistical Machine Translation (SMT). However, most existing parallel c…
Boundary DetectionDomain AdaptationInformation RetrievalMachine Translation+2Bidirectional Chinese and English Passive Sentences Dataset for Machine Translation
Machine Translation (MT) evaluation has gone beyond metrics, towards more specific linguistic phenomena. Regarding English-Chinese language pairs, passive sentences are constructed and distributed differently due to lang…
Machine TranslationUniversal Semantic Tagging for English and Mandarin Chinese
Universal Semantic Tagging aims to provide lightweight unified analysis for all languages at the word level. Though the proposed annotation scheme is conceptually promising, the feasibility is only examined in four Indo{…
ASPEC: Asian Scientific Paper Excerpt Corpus
In this paper, we describe the details of the ASPEC (Asian Scientific Paper Excerpt Corpus), which is the first large-size parallel corpus of scientific paper domain. ASPEC was constructed in the Japanese-Chinese machine…
Machine TranslationTranslation