paper-with-me

홈 › Papers

Building the Language Resource for a Cebuano-Filipino Neural Machine Translation System

2021-10-05 · Kristine Mae Adlaon, Nelson Marcos

Parallel corpus is a critical resource in machine learning-based translation. The task of collecting, extracting, and aligning texts in order to build an acceptable corpus for doing the translation is very tedious most especially for low-resource languages. In this paper, we present the efforts made to build a parallel corpus for Cebuano and Filipino from two different domains: biblical texts and the web. For the biblical resource, subword unit translation for verbs and copy-able approach for nouns were applied to correct inconsistencies in the translation. This correction mechanism was applied as a preprocessing technique. On the other hand, for Wikipedia being the main web resource, commonly occurring topic segments were extracted from both the source and the target languages. These observed topic segments are unique in 4 different categories. The identification of these topic segments may be used for the automatic extraction of sentences. A Recurrent Neural Network was used to implement the translation using OpenNMT sequence modeling tool in TensorFlow. The two different corpora were then evaluated by using them as two separate inputs in the neural network. Results have shown a difference in BLEU scores in both corpora.

📄 PDF Abstract BibTeX arXiv:2110.15716

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Exploring Word Alignment towards an Efficient Sentence Aligner for Filipino and Cebuano Languages

2022-10-01 · loresmt (COLING) 2022 10 · Jenn Leana Fernandez, Kristine Mae M. Adlaon

Building a robust machine translation (MT) system requires a large amount of parallel corpus which is an expensive resource for low-resourced languages. The two major languages being spoken in the Philippines which are F…

Machine TranslationSentenceTranslationWord Alignment

A Baseline Readability Model for Cebuano

2022-03-31 · NAACL (BEA) 2022 7 · Lloyd Lois Antonie Reyes, Michael Antonio Ibañez, Ranz Sapinit, Mohammed Hussien 외

In this study, we developed the first baseline readability model for the Cebuano language. Cebuano is the second most-used native language in the Philippines with about 27.5 million speakers. As the baseline, we extracte…

model

FilBench: Can LLMs Understand and Generate Filipino?

2025-08-05 · Lester James V. Miranda, Elyanah Aco, Conner Manuel, Jan Christian Blaise Cruz 외 arxiv

Despite the impressive performance of LLMs on English-based tasks, little is known about their capabilities in specific languages such as Filipino. In this work, we address this gap by introducing FilBench, a Filipino-ce…

Reading Comprehension

Towards Automatic Construction of Filipino WordNet: Word Sense Induction and Synset Induction Using Sentence Embeddings

2022-04-07 · Dan John Velasco, Axel Alba, Trisha Gail Pelagio, Bryce Anthony Ramirez 외

Wordnets are indispensable tools for various natural language processing applications. Unfortunately, wordnets get outdated, and producing or updating wordnets can be slow and costly in terms of time and resources. This …

Language ModelingLanguage ModellingSentenceSentence Embeddings+2

Pagsusuri ng RNN-based Transfer Learning Technique sa Low-Resource Language

2020-10-13 · Dan John Velasco

Low-resource languages such as Filipino suffer from data scarcity which makes it challenging to develop NLP applications for Filipino language. The use of Transfer Learning (TL) techniques alleviates this problem in low-…

Language ModelingLanguage ModellingTransfer Learning