Automatic Translation of Scientific Documents in the HAL Archive
This paper describes the development of a statistical machine translation system between French and English for scientific papers. This system will be closely integrated into the French HAL open archive, a collection of more than 100.000 scientific papers. We describe the creation of in-domain parallel and monolingual corpora, the development of a domain specific translation system with the created resources, and its adaptation using monolingual resources only. These techniques allowed us to improve a generic system by more than 10 BLEU points.
Code (0)
등록된 구현이 없습니다.
Tasks
Domain AdaptationMachine TranslationTranslationSimilar Papers 제목 키워드 기반
Identifying Documents In-Scope of a Collection from Web Archives
Web archive data usually contains high-quality documents that are very useful for creating specialized collections of documents, e.g., scientific digital libraries and repositories of technical reports. In doing so, ther…
ATEM: A Topic Evolution Model for the Detection of Emerging Topics in Scientific Archives
This paper presents ATEM, a novel framework for studying topic evolution in scientific archives. ATEM is based on dynamic topic modeling and dynamic graph embedding techniques that explore the dynamics of content and cit…
ArticlesDynamic graph embeddingDynamic Topic ModelingGraph EmbeddingThe Scielo Corpus: a Parallel Corpus of Scientific Publications for Biomedicine
The biomedical scientific literature is a rich source of information not only in the English language, for which it is more abundant, but also in other languages, such as Portuguese, Spanish and French. We present the fi…
Machine TranslationTranslationHandwriting Classification for the Analysis of Art-Historical Documents
Digitized archives contain and preserve the knowledge of generations of scholars in millions of documents. The size of these archives calls for automatic analysis since a manual analysis by specialists is often too expen…
ClassificationGeneral ClassificationOptical Character Recognition (OCR)text-classification+1Translation Using JAPIO Patent Corpora: JAPIO at WAT2016
We participate in scientific paper subtask (ASPEC-EJ/CJ) and patent subtask (JPC-EJ/CJ/KJ) with phrase-based SMT systems which are trained with its own patent corpora. Using larger corpora than those prepared by the work…
Information RetrievalMachine TranslationTranslation