Introducing the Asian Language Treebank (ALT)
This paper introduces the ALT project initiated by the Advanced Speech Translation Research and Development Promotion Center (ASTREC), NICT, Kyoto, Japan. The aim of this project is to accelerate NLP research for Asian languages such as Indonesian, Japanese, Khmer, Laos, Malay, Myanmar, Philippine, Thai and Vietnamese. The original resource for this project was English articles that were randomly selected from Wikinews. The project has so far created a corpus for Myanmar and will extend in scope to include other languages in the near future. A 20000-sentence corpus of Myanmar that has been manually translated from an English corpus has been word segmented, word aligned, part-of-speech tagged and constituency parsed by human annotators. In this paper, we present the implementation steps for creating the treebank in detail, including a description of the ALT web-based treebanking tool. Moreover, we report statistics on the annotation quality of the Myanmar treebank created so far.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesSentenceTranslationSimilar Papers 제목 키워드 기반
First Steps towards Universal Dependencies for Laz
This paper presents the first treebank for the Laz language, which is also the first Universal Dependencies Treebank for a South Caucasian language. This treebank aims to create a syntactically and morphologically annota…
Dependency ParsingSimilar Southeast Asian Languages: Corpus-Based Case Study on Thai-Laotian and Malay-Indonesian
This paper illustrates the similarity between Thai and Laotian, and between Malay and Indonesian, based on an investigation on raw parallel data from Asian Language Treebank. The cross-lingual similarity is investigated …
Machine TranslationTranslationWord AlignmentAn Overview of BPPT's Indonesian Language Resources
This paper describes various Indonesian language resources that Agency for the Assessment and Application of Technology (BPPT) has developed and collected since mid 80{'}s when we joined MMTS (Multilingual Machine Transl…
Machine Translationspeech-recognitionSpeech RecognitionSpeech Synthesis+1Treebank Embedding Vectors for Out-of-domain Dependency Parsing
A recent advance in monolingual dependency parsing is the idea of a treebank embedding vector, which allows all treebanks for a particular language to be used as training data while at the same time allowing the model to…
Dependency ParsingIntroducing YakuToolkit. Yakut Treebank and Morphological Analyzer.
This poster presents the first publicly available treebank of Yakut, a Turkic language spoken in Russia, and a morphological analyzer for this language. The treebank was annotated following the Universal Dependencies (UD…