paper-with-me

홈 › Papers

Demo of Sanskrit-Hindi SMT System

2018-04-13 · Rajneesh Pandey, Atul Kr. Ojha, Girish Nath Jha

The demo proposal presents a Phrase-based Sanskrit-Hindi (SaHiT) Statistical Machine Translation system. The system has been developed on Moses. 43k sentences of Sanskrit-Hindi parallel corpus and 56k sentences of a monolingual corpus in the target language (Hindi) have been used. This system gives 57 BLEU score.

📄 PDF Abstract BibTeX arXiv:1804.06716

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

A SANSKRIT TO HINDI LANGUAGE MACHINE TRANSLATOR USING RULE BASED APPROACH

2020-12-01 · ICON 2020 12 · Prateek Agrawal, Vishu Madaan

Hindi and Sanskrit both the languages are having the same script i.e. Devnagari Script which results in few basic similarities in their grammar rules. As we know that Hindi ranks fourth in terms of speaker’s size in the …

POSSentence

An Augmented Translation Technique for low Resource language pair: Sanskrit to Hindi translation

2020-06-09 · Rashi Kumar, Piyush Jha, Vineet Sahula

Neural Machine Translation (NMT) is an ongoing technique for Machine Translation (MT) using enormous artificial neural network. It has exhibited promising outcomes and has shown incredible potential in solving challengin…

Dimensionality ReductionMachine TranslationNMTTranslation

Samasāmayik: A Parallel Dataset for Hindi-Sanskrit Machine Translation

2026-03-25 · N J Karthika, Keerthana Suryanarayanan, Jahanvi Purohit, Ganesh Ramakrishnan 외 arxiv

We release Samasāmayik, a novel, meticulously curated, large-scale Hindi-Sanskrit corpus, comprising 92,196 parallel sentences. Unlike most data available in Sanskrit, which focuses on classical era text and poetry, this…

Machine Translation

Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece

2026-01-05 · Anshul Kumar arxiv

Tokens are the basic units of Large Language Models (LLMs). LLMs rely on tokenizers to segment text into these tokens, and tokenization is the primary determinant of computational and inference cost. Sanskrit, one of the…

Attention based Sequence to Sequence Learning for Machine Translation of Low Resourced Indic Languages -- A case of Sanskrit to Hindi

2021-09-07 · Vishvajit Bakarola, Jitendra Nasriwala

Deep Learning techniques are powerful in mimicking humans in a particular set of problems. They have achieved a remarkable performance in complex learning tasks. Deep learning inspired Neural Machine Translation (NMT) is…

Machine TranslationNMTTranslation