A Distributed Resource Repository for Cloud-Based Machine Translation
In this paper, we present the architecture of a distributed resource repository developed for collecting training data for building customized statistical machine translation systems. The repository is designed for the cloud-based translation service integrated in the Let'sMT! platform which is about to be launched to the public. The system includes important features such as automatic import and alignment of textual documents in a variety of formats, a flexible database for meta-information using modern key-value stores and a grid-based backend for running off-line processes. The entire system is very modular and supports highly distributed setups to enable a maximum of flexibility and scalability. The system uses secure connections and includes an effective permission management to ensure data integrity. In this paper, we also take a closer look at the task of sentence alignment. The process of alignment is extremely important for the success of translation models trained on the platform. Alignment decisions significantly influence the quality of SMT engines.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationManagementSentenceTranslationSimilar Papers 제목 키워드 기반
The OPUS Resource Repository: An Open Package for Creating Parallel Corpora and Machine Translation Services
This paper presents a flexible and powerful system for creating parallel corpora and for running neural machine translation services. Our package provides a scalable data repository backend that offers transparent data p…
Machine TranslationTranslationEfficient Resource Scheduling for Distributed Infrastructures Using Negotiation Capabilities
In the past few decades, the rapid development of information and internet technologies has spawned massive amounts of data and information. The information explosion drives many enterprises or individuals to seek to ren…
Cloud ComputingSchedulingOPUS-MT – Building open translation services for the World
This paper presents OPUS-MT a project that focuses on the development of free resources and tools for machine translation. The current status is a repository of over 1,000 pre-trained neural machine translation models th…
Machine TranslationTranslationEnhancing Assamese NLP Capabilities: Introducing a Centralized Dataset Repository
This paper introduces a centralized, open-source dataset repository designed to advance NLP and NMT for Assamese, a low-resource language. The repository, available at GitHub, supports various tasks like sentiment analys…
DiversityMachine Translationnamed-entity-recognitionNamed Entity Recognition+4A Survey on Machine Learning for Geo-Distributed Cloud Data Center Management
Cloud workloads today are typically managed in a distributed environment and processed across geographically distributed data centers. Cloud service providers have been distributing data centers globally to reduce operat…
BIG-bench Machine LearningManagement