paper-with-me

Papers

Improving Indonesian Text Classification Using Multilingual Language Model

2020-09-12 · Ilham Firdausi Putra, Ayu Purwarianti

Compared to English, the amount of labeled data for Indonesian text classification tasks is very small. Recently developed multilingual language models have shown its ability to create multilingual representations effectively. This paper investigates the effect of combining English and Indonesian data on building Indonesian text classification (e.g., sentiment analysis and hate speech) using multilingual language models. Using the feature-based approach, we observe its performance on various data sizes and total added English data. The experiment showed that the addition of English data, especially if the amount of Indonesian data is small, improves performance. Using the fine-tuning approach, we further showed its effectiveness in utilizing the English language to build Indonesian text classification models.

📄 PDF Abstract BibTeX arXiv:2009.05713

Code (1)

ilhamfp/indonesian-text-classification-multilingual 공식 구현 pytorch

Tasks

ClassificationGeneral ClassificationLanguage ModelingLanguage ModellingmodelSentiment Analysistext-classificationText Classification

Similar Papers 제목 키워드 기반

Indonesian-English Code-Switching Speech Synthesizer Utilizing Multilingual STEN-TTS and Bert LID

2024-12-26 · Ahmad Alfani Handoyo, Chung Tran, Dessi Puji Lestari, Sakriani Sakti

Multilingual text-to-speech systems convert text into speech across multiple languages. In many cases, text sentences may contain segments in different languages, a phenomenon known as code-switching. This is particularl…

Language Identificationtext-to-speechText to Speech

idT5: Indonesian Version of Multilingual T5 Transformer

2023-02-02 · Mukhlish Fuadi, Adhi Dharma Wibawa, Surya Sumpeno

Indonesian language is spoken by almost 200 million people and is the 10th most spoken language in the world, but it is under-represented in NLP (Natural Language Processing) research. A sparsity of language resources ha…

Question AnsweringQuestion GenerationQuestion-GenerationSentiment Analysis

Fine-tuning Pretrained Multilingual BERT Model for Indonesian Aspect-based Sentiment Analysis

2021-03-05 · Annisa Nurul Azhar, Masayu Leylia Khodra

Although previous research on Aspect-based Sentiment Analysis (ABSA) for Indonesian reviews in hotel domain has been conducted using CNN and XGBoost, its model did not generalize well in test data and high number of OOV …

Aspect-Based Sentiment AnalysisAspect-Based Sentiment Analysis (ABSA)Sentiment Analysis

Does Language Shift Break Medical Vision-Language Models? Indonesian Radiology Visual Question Answering Case Study

2026-06-02 · Pieter Christy Yan Yudhistira, Dzaki Rafif Malik, Novanto Yudistira arxiv

Medical Vision-Language Models (VLMs) are typically evaluated on English radiology visual question answering benchmarks, leaving their robustness under non-English clinical language largely unexplored. We introduce IndoR…

Visual Question Answering

COPAL-ID: Indonesian Language Reasoning with Local Culture and Nuances

2023-11-02 · Haryo Akbarianto Wibowo, Erland Hilman Fuadi, Made Nindyatama Nityasya, Radityo Eko Prasojo 외

We present COPAL-ID, a novel, public Indonesian language common sense reasoning dataset. Unlike the previous Indonesian COPA dataset (XCOPA-ID), COPAL-ID incorporates Indonesian local and cultural nuances, and therefore,…

Common Sense Reasoning