paper-with-me

Papers

zNLP: Identifying Parallel Sentences in Chinese-English Comparable Corpora

2017-08-01 · WS 2017 8 · Zheng Zhang, Pierre Zweigenbaum

This paper describes the zNLP system for the BUCC 2017 shared task. Our system identifies parallel sentence pairs in Chinese-English comparable corpora by translating word-by-word Chinese sentences into English, using the search engine Solr to select near-parallel sentences and then by using an SVM classifier to identify true parallel sentences from the previous results. It obtains an F1-score of 45{\%} (resp. 32{\%}) on the test (training) set.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentence

Methods 이 논문이 사용한 방법론

SVM A Support Vector Machine, or SVM, is a non-parametric supervised learning model. For non-linear classification and regression, they utilise the kernel trick to map inputs…

Similar Papers 제목 키워드 기반

Identify Bilingual Patterns and Phrases from a Bilingual Sentence Pair

2021-10-01 · ROCLING 2021 10 · Yi-Jyun Chen, Hsin-Yun Chung, Jason S. Chang

This paper presents a method for automatically identifying bilingual grammar patterns and extracting bilingual phrase instances from a given English-Chinese sentence pair. In our approach, the English-Chinese sentence pa…

Machine TranslationSentenceTranslation

UM-Corpus: A Large English-Chinese Parallel Corpus for Statistical Machine Translation

2014-05-01 · LREC 2014 5 · Liang Tian, Derek F. Wong, Lidia S. Chao, Paulo Quaresma 외

Parallel corpus is a valuable resource for cross-language information retrieval and data-driven natural language processing systems, especially for Statistical Machine Translation (SMT). However, most existing parallel c…

Boundary DetectionDomain AdaptationInformation RetrievalMachine Translation+2

Bidirectional Chinese and English Passive Sentences Dataset for Machine Translation

2026-03-16 · Xinyue Ma, Pol Pastells, Mireia Farrús, Mariona Taulé arxiv

Machine Translation (MT) evaluation has gone beyond metrics, towards more specific linguistic phenomena. Regarding English-Chinese language pairs, passive sentences are constructed and distributed differently due to lang…

Machine Translation

Universal Semantic Tagging for English and Mandarin Chinese

2021-06-01 · NAACL 2021 4 · Wenxi Li, Yiyang Hou, Yajie Ye, Li Liang 외

Universal Semantic Tagging aims to provide lightweight unified analysis for all languages at the word level. Though the proposed annotation scheme is conceptually promising, the feasibility is only examined in four Indo{…

ASPEC: Asian Scientific Paper Excerpt Corpus

2016-05-01 · LREC 2016 5 · Toshiaki Nakazawa, Manabu Yaguchi, Kiyotaka Uchimoto, Masao Utiyama 외

In this paper, we describe the details of the ASPEC (Asian Scientific Paper Excerpt Corpus), which is the first large-size parallel corpus of scientific paper domain. ASPEC was constructed in the Japanese-Chinese machine…

Machine TranslationTranslation