paper-with-me

Papers

Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry

2025-05-30 · Sujeet Kumar, Pretam Ray, Abhinay Beerukuri, Shrey Kamoji, Manoj Balaji Jagadeeshan, Pawan Goyal

Sanskrit, an ancient language with a rich linguistic heritage, presents unique challenges for automatic speech recognition (ASR) due to its phonemic complexity and the phonetic transformations that occur at word junctures, similar to the connected speech found in natural conversations. Due to these complexities, there has been limited exploration of ASR in Sanskrit, particularly in the context of its poetic verses, which are characterized by intricate prosodic and rhythmic patterns. This gap in research raises the question: How can we develop an effective ASR system for Sanskrit, particularly one that captures the nuanced features of its poetic form? In this study, we introduce Vedavani, the first comprehensive ASR study focused on Sanskrit Vedic poetry. We present a 54-hour Sanskrit ASR dataset, consisting of 30,779 labelled audio samples from the Rig Veda and Atharva Veda. This dataset captures the precise prosodic and rhythmic features that define the language. We also benchmark the dataset on various state-of-the-art multilingual speech models.$^{1}$ Experimentation revealed that IndicWhisper performed the best among the SOTA models.

📄 PDF Abstract BibTeX arXiv:2506.00145

Code (1)

sujeetnlp/vedavani 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

The Treebank of Vedic Sanskrit

2020-05-01 · LREC 2020 5 · Oliver Hellwig, Salvatore Scarlata, Elia Ackermann, Paul Widmer

This paper introduces the first treebank of Vedic Sanskrit, a morphologically rich ancient Indian language that is of central importance for linguistic and historical research. The selection of the more than 3,700 senten…

Detecting Diachronic Syntactic Developments in Presence of Bias Terms

2022-06-01 · LT4HALA (LREC) 2022 6 · Oliver Hellwig, Sven Sellmer

Corpus-based studies of diachronic syntactic changes are typically guided by the results of previous qualitative research. When such results are missing or, as is the case for Vedic Sanskrit, are restricted to small part…

Accent Placement Models for Rigvedic Sanskrit Text

2025-11-28 · Akhil Rajeev P, Annarao Kulkarni arxiv

The Rigveda, among the oldest Indian texts in Vedic Sanskrit, employs a distinctive pitch-accent system : udātta, anudātta, svarita whose marks encode melodic and interpretive cues but are often absent from modern e-text…

parameter-efficient fine-tuning

Samasāmayik: A Parallel Dataset for Hindi-Sanskrit Machine Translation

2026-03-25 · N J Karthika, Keerthana Suryanarayanan, Jahanvi Purohit, Ganesh Ramakrishnan 외 arxiv

We release Samasāmayik, a novel, meticulously curated, large-scale Hindi-Sanskrit corpus, comprising 92,196 parallel sentences. Unlike most data available in Sanskrit, which focuses on classical era text and poetry, this…

Machine Translation

Aesthetics of Sanskrit Poetry from the Perspective of Computational Linguistics: A Case Study Analysis on Siksastaka

2023-08-14 · Jivnesh Sandhan, Amruta Barbadikar, Malay Maity, Pavankumar Satuluri 외

Sanskrit poetry has played a significant role in shaping the literary and cultural landscape of the Indian subcontinent for centuries. However, not much attention has been devoted to uncovering the hidden beauty of Sansk…