paper-with-me

홈 › Papers

Sāmayik: A Benchmark and Dataset for English-Sanskrit Translation

2023-05-23 · Ayush Maheshwari, Ashim Gupta, Amrith Krishna, Atul Kumar Singh, Ganesh Ramakrishnan, G. Anil Kumar, Jitin Singla

We release S\={a}mayik, a dataset of around 53,000 parallel English-Sanskrit sentences, written in contemporary prose. Sanskrit is a classical language still in sustenance and has a rich documented heritage. However, due to the limited availability of digitized content, it still remains a low-resource language. Existing Sanskrit corpora, whether monolingual or bilingual, have predominantly focused on poetry and offer limited coverage of contemporary written materials. S\={a}mayik is curated from a diverse range of domains, including language instruction material, textual teaching pedagogy, and online tutorials, among others. It stands out as a unique resource that specifically caters to the contemporary usage of Sanskrit, with a primary emphasis on prose writing. Translation models trained on our dataset demonstrate statistically significant improvements when translating out-of-domain contemporary corpora, outperforming models trained on older classical-era poetry datasets. Finally, we also release benchmark models by adapting four multilingual pre-trained models, three of them have not been previously exposed to Sanskrit for translating between English and Sanskrit while one of them is multi-lingual pre-trained translation model including English and Sanskrit. The dataset and source code is present at https://github.com/ayushbits/saamayik.

📄 PDF Abstract BibTeX arXiv:2305.14004

Code (1)

ayushbits/saamayik 공식 구현 pytorch

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Samasāmayik: A Parallel Dataset for Hindi-Sanskrit Machine Translation

2026-03-25 · N J Karthika, Keerthana Suryanarayanan, Jahanvi Purohit, Ganesh Ramakrishnan 외 arxiv

We release Samasāmayik, a novel, meticulously curated, large-scale Hindi-Sanskrit corpus, comprising 92,196 parallel sentences. Unlike most data available in Sanskrit, which focuses on classical era text and poetry, this…

Machine Translation

Anveshana: A New Benchmark Dataset for Cross-Lingual Information Retrieval On English Queries and Sanskrit Documents

2025-05-26 · Manoj Balaji Jagadeeshan, Prince Raj, Pawan Goyal

The study presents a comprehensive benchmark for retrieving Sanskrit documents using English queries, focusing on the chapters of the Srimadbhagavatam. It employs a tripartite approach: Direct Retrieval (DR), Translation…

Cross-Lingual Information RetrievalInformation RetrievalRAGRetrieval+1

Itihasa: A large-scale corpus for Sanskrit to English translation

2021-06-06 · ACL (WAT) 2021 8 · Rahul Aralikatte, Miryam de Lhoneux, Anoop Kunchukuttan, Anders Søgaard

This work introduces Itihasa, a large-scale translation dataset containing 93,000 pairs of Sanskrit shlokas and their English translations. The shlokas are extracted from two Indian epics viz., The Ramayana and The Mahab…

Machine TranslationTranslation

Improving Neural Machine Translation for Sanskrit-English

2020-12-01 · ICON 2020 12 · Ravneet Punia, Aditya Sharma, Sarthak Pruthi, Minni Jain

Sanskrit is one of the oldest languages of the Asian Subcontinent that fell out of common usage around 600 B.C. In this paper, we attempt to translate Sanskrit to English using Neural Machine Translation approaches based…

Machine Translationreinforcement-learningReinforcement Learning (RL)Transfer Learning+1

An Augmented Translation Technique for low Resource language pair: Sanskrit to Hindi translation

2020-06-09 · Rashi Kumar, Piyush Jha, Vineet Sahula

Neural Machine Translation (NMT) is an ongoing technique for Machine Translation (MT) using enormous artificial neural network. It has exhibited promising outcomes and has shown incredible potential in solving challengin…

Dimensionality ReductionMachine TranslationNMTTranslation