paper-with-me

Papers

Multilingual TinyStories: A Synthetic Combinatorial Corpus of Indic Children's Stories for Training Small Language Models

2026-03-15 · Deepon Halder, Angira Mukherjee arxiv

The development of robust language models for low-resource languages is frequently bottlenecked by the scarcity of high-quality, coherent, and domain-appropriate training corpora. In this paper, we introduce the Multilingual TinyStories dataset, a large-scale, synthetically generated collection of children's stories encompassing 17 Indian languages. Designed specifically for the training and evaluation of Small Language Models (SLMs), the corpus provides simple, narrative-driven text strictly localized to native scripts. We detail our hybrid curation pipeline, which leverages the Sarvam-M language model and a novel combinatorial prompt engineering framework for native generation, coupled with the Google Translate API for large-scale cross-lingual expansion. Through strict programmatic filtering, we compiled 132,942 stories and over 93.9 million tokens in our release, serving as a foundational resource for multilingual language modeling and transfer learning in the Indic linguistic sphere.

📄 PDF Abstract BibTeX arXiv:2603.14563

Code (0)

등록된 구현이 없습니다.

Tasks

Prompt EngineeringTransfer Learning

Similar Papers 제목 키워드 기반

NIT Rourkela Machine Translation(MT) System Submission to WAT 2022 for MultiIndicMT: An Indic Language Multilingual Shared Task

2022-10-01 · WAT 2022 10 · Sudhansu Bala Das, Atharv Biradar, Tapas Kumar Mishra, Bidyut Kumar Patra

Multilingual Neural Machine Translation (MNMT) exhibits incredible performance with the development of a single translation model for many languages. Previous studies on multilingual translation reveal that multilingual …

DecoderMachine TranslationNMTTranslation

Cross-Lingual Interleaving for Speech Language Models

2025-12-01 · Adel Moumen, Guangzhi Sun, Philip C. Woodland arxiv

Spoken Language Models (SLMs) aim to learn linguistic competence directly from speech using discrete units, widening access to Natural Language Processing (NLP) technologies for languages with limited written resources. …

BERTtime Stories: Investigating the Role of Synthetic Story Data in Language pre-training

2024-10-20 · Nikitas Theodoropoulos, Giorgos Filandrianos, Vassilis Lyberatos, Maria Lymperaiou 외

We describe our contribution to the Strict and Strict-Small tracks of the 2nd iteration of the BabyLM Challenge. The shared task is centered around efficient pre-training given data constraints motivated by human develop…

Language ModelingLanguage Modelling

Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training

2025-04-02 · Zhijun Wang, Jiahuan Li, Hao Zhou, Rongxiang Weng 외

Large language models (LLMs) exhibit remarkable multilingual capabilities despite the extreme language imbalance in the pre-training data. In this paper, we closely examine the reasons behind this phenomenon, focusing on…

Language ModelingLanguage Modelling

IITP-MT at WAT2021: Indic-English Multilingual Neural Machine Translation using Romanized Vocabulary

2021-08-01 · ACL (WAT) 2021 8 · Ramakrishna Appicharla, Kamal Kumar Gupta, Asif Ekbal, Pushpak Bhattacharyya

This paper describes the systems submitted to WAT 2021 MultiIndicMT shared task by IITP-MT team. We submit two multilingual Neural Machine Translation (NMT) systems (Indic-to-English and English-to-Indic). We romanize al…

Machine TranslationNMTTranslation