paper-with-me

홈 › Papers

BanglaSTEM: A Parallel Corpus for Technical Domain Bangla-English Translation

2025-11-05 · Kazi Reyazul Hasan, Mubasshira Musarrat, A. B. M. Alim Al Islam, Muhammad Abdullah Adnan arxiv

Large language models work well for technical problem solving in English but perform poorly when the same questions are asked in Bangla. A simple solution would be to translate Bangla questions into English first and then use these models. However, existing Bangla-English translation systems struggle with technical terms. They often mistranslate specialized vocabulary, which changes the meaning of the problem and leads to wrong answers. We present BanglaSTEM, a dataset of 5,000 carefully selected Bangla-English sentence pairs from STEM fields including computer science, mathematics, physics, chemistry, and biology. We generated over 12,000 translations using language models and then used human evaluators to select the highest quality pairs that preserve technical terminology correctly. We train a T5-based translation model on BanglaSTEM and test it on two tasks: generating code and solving math problems. Our results show significant improvements in translation accuracy for technical content, making it easier for Bangla speakers to use English-focused language models effectively. Both the BanglaSTEM dataset and the trained translation model are publicly released at https://huggingface.co/reyazul/BanglaSTEM-T5.

📄 PDF Abstract BibTeX arXiv:2511.03498

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BanglaNirTox: A Large-scale Parallel Corpus for Explainable AI in Bengali Text Detoxification

2025-11-03 · Ayesha Afroza Mohsin, Mashrur Ahsan, Nafisa Maliyat, Shanta Maria 외 arxiv

Toxic language in Bengali remains prevalent, especially in online environments, with few effective precautions against it. Although text detoxification has seen progress in high-resource languages, Bengali remains undere…

Investigating self-supervised, weakly supervised and fully supervised training approaches for multi-domain automatic speech recognition: a study on Bangladeshi Bangla

2022-10-24 · Ahnaf Mozib Samin, M. Humayon Kobir, Md. Mushtaq Shahriyar Rafee, M. Firoz Ahmed 외

Despite huge improvements in automatic speech recognition (ASR) employing neural networks, ASR systems still suffer from a lack of robustness and generalizability issues due to domain shifting. This is mainly because pri…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

The EuroPat Corpus: A Parallel Corpus of European Patent Data

2022-06-01 · LREC 2022 6 · Kenneth Heafield, Elaine Farrow, Jelmer Van der Linde, Gema Ramírez-Sánchez 외

We present the EuroPat corpus of patent-specific parallel data for 6 official European languages paired with English: German, Spanish, French, Croatian, Norwegian, and Polish. The filtered parallel corpora range in size …

Machine TranslationTranslation

The LTRC Hindi-Telugu Parallel Corpus

2022-06-01 · LREC 2022 6 · Vandan Mujadia, Dipti Sharma

We present the Hindi-Telugu Parallel Corpus of different technical domains such as Natural Science, Computer Science, Law and Healthcare along with the General domain. The qualitative corpus consists of 700K parallel sen…

DiversityMachine TranslationTranslation

Oral to Web: Digitizing 'Zero Resource'Languages of Bangladesh

2026-03-05 · Mohammad Mamun Or Rashid arxiv

We present the Multilingual Cloud Corpus, the first national-scale, parallel, multimodal linguistic dataset of Bangladesh's ethnic and indigenous languages. Despite being home to approximately 40 minority languages spann…