An Overview of BPPT's Indonesian Language Resources
This paper describes various Indonesian language resources that Agency for the Assessment and Application of Technology (BPPT) has developed and collected since mid 80{'}s when we joined MMTS (Multilingual Machine Translation System), an international project coordinated by CICC-Japan to develop a machine translation system for five Asian languages (Bahasa Indonesia, Malay, Thai, Japanese, and Chinese). Since then, we have been actively doing many types of research in the field of statistical machine translation, speech recognition, and speech synthesis which requires many text and speech corpus. Most recent cooperation within ASEAN-IVO is the development of Indonesian ALT (Asian Language Treebank) has added new NLP tools.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine Translationspeech-recognitionSpeech RecognitionSpeech SynthesisTranslationSimilar Papers 제목 키워드 기반
Scoping natural language processing in Indonesian and Malay for education applications
Indonesian and Malay are underrepresented in the development of natural language processing (NLP) technologies and available resources are difficult to find. A clear picture of existing work can invigorate and inform how…
Reading ComprehensionSentiment AnalysisNusaCrowd: A Call for Open and Reproducible NLP Research in Indonesian Languages
At the center of the underlying issues that halt Indonesian natural language processing (NLP) research advancement, we find data scarcity. Resources in Indonesian languages, especially the local ones, are extremely scarc…
IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP
Although the Indonesian language is spoken by almost 200 million people and the 10th most spoken language in the world, it is under-represented in NLP research. Previous work on Indonesian has been hampered by a lack of …
BenchmarkingLanguage ModelingLanguage ModellingOne Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia
NLP research is impeded by a lack of resources and awareness of the challenges presented by underrepresented languages and dialects. Focusing on the languages spoken in Indonesia, the second most linguistically diverse a…
DriveThru: a Document Extraction Platform and Benchmark Datasets for Indonesian Local Language Archives
Indonesia is one of the most diverse countries linguistically. However, despite this linguistic diversity, Indonesian languages remain underrepresented in Natural Language Processing (NLP) research and technologies. In t…
Optical Character RecognitionOptical Character Recognition (OCR)