Handling Indonesian Clitics: A Dataset Comparison for an Indonesian-English Statistical Machine Translation System
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationTranslationWord AlignmentSimilar Papers 제목 키워드 기반
IndoPref: A Multi-Domain Pairwise Preference Dataset for Indonesian
Over 200 million people speak Indonesian, yet the language remains significantly underrepresented in preference-based research for large language models (LLMs). Most existing multilingual datasets are derived from Englis…
Indonesian-English Code-Switching Speech Synthesizer Utilizing Multilingual STEN-TTS and Bert LID
Multilingual text-to-speech systems convert text into speech across multiple languages. In many cases, text sentences may contain segments in different languages, a phenomenon known as code-switching. This is particularl…
Language Identificationtext-to-speechText to SpeechLocation-based Twitter Filtering for the Creation of Low-Resource Language Datasets in Indonesian Local Languages
Twitter contains an abundance of linguistic data from the real world. We examine Twitter for user-generated content in low-resource languages such as local Indonesian. For NLP to work in Indonesian, it must consider loca…
Cultural Vocal Bursts Intensity PredictionNusaCrowd: A Call for Open and Reproducible NLP Research in Indonesian Languages
At the center of the underlying issues that halt Indonesian natural language processing (NLP) research advancement, we find data scarcity. Resources in Indonesian languages, especially the local ones, are extremely scarc…
COPAL-ID: Indonesian Language Reasoning with Local Culture and Nuances
We present COPAL-ID, a novel, public Indonesian language common sense reasoning dataset. Unlike the previous Indonesian COPA dataset (XCOPA-ID), COPAL-ID incorporates Indonesian local and cultural nuances, and therefore,…
Common Sense Reasoning