paper-with-me

Papers

Accent Placement Models for Rigvedic Sanskrit Text

2025-11-28 · Akhil Rajeev P, Annarao Kulkarni arxiv

The Rigveda, among the oldest Indian texts in Vedic Sanskrit, employs a distinctive pitch-accent system : udātta, anudātta, svarita whose marks encode melodic and interpretive cues but are often absent from modern e-texts. This work develops a parallel corpus of accented-unaccented ślokas and conducts a controlled comparison of three strategies for automatic accent placement in Rigvedic verse: (i) full fine-tuning of ByT5, a byte-level Transformer that operates directly on Unicode combining marks, (ii) a from-scratch BiLSTM-CRF sequence-labeling baseline, and (iii) LoRA-based parameter-efficient fine-tuning atop ByT5. Evaluation uses Word Error Rate (WER) and Character Error Rate (CER) for orthographic fidelity, plus a task-specific Diacritic Error Rate (DER) that isolates accent edits. Full ByT5 fine-tuning attains the lowest error across all metrics; LoRA offers strong efficiency-accuracy trade-offs, and BiLSTM-CRF serves as a transparent baseline. The study underscores practical requirements for accent restoration - Unicode-safe preprocessing, mark-aware tokenization, and evaluation that separates grapheme from accent errors - and positions heritage-language technology as an emerging NLP area connecting computational modeling with philological and pedagogical aims. Results establish reproducible baselines for Rigvedic accent restoration and provide guidance for downstream tasks such as accent-aware OCR, ASR/chant synthesis, and digital scholarship.

📄 PDF Abstract BibTeX arXiv:2511.23088

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuning

Similar Papers 제목 키워드 기반

Web based System for Derivational Process of Kṛdanta based on Pāṇinian Grammatical Tradition

2022-06-01 · WILDRE (LREC) 2022 6 · Sumit Sharma, Subhash Chandra

Each text of the Sanskrit literature is wadded with the uses of Sanskrit kṛdanta (participles). The knowledge and formation process of Sanskrit kṛdanta play a key role to understand the meaning of a particular kṛdanta wo…

San-BERT: Extractive Summarization for Sanskrit Documents using BERT and it's variants

2023-04-04 · Kartik Bhatnagar, Sampath Lonka, Jammi Kunal, Mahabala Rao M G

In this work, we develop language models for the Sanskrit language, namely Bidirectional Encoder Representations from Transformers (BERT) and its variants: A Lite BERT (ALBERT), and Robustly Optimized BERT (RoBERTa) usin…

ClusteringExtractive SummarizationExtractive Text SummarizationText Summarization

Embeddings models for Buddhist Sanskrit

2022-06-01 · LREC 2022 6 · Ligeia Lugli, Matej Martinc, Andraž Pelicon, Senja Pollak

The paper presents novel resources and experiments for Buddhist Sanskrit, broadly defined here including all the varieties of Sanskrit in which Buddhist texts have been transmitted. We release a novel corpus of Buddhist …

Semantic SimilaritySemantic Textual SimilarityTransfer LearningWord Similarity

A Benchmark Corpus and Neural Approach for Sanskrit Derivative Nouns Analysis

2020-10-24 · Arun Kumar Singh, Sushant Dave, Dr. Prathosh A. P., Prof. Brejesh Lall 외

This paper presents first benchmark corpus of Sanskrit Pratyaya (suffix) and inflectional words (padas) formed due to suffixes along with neural network based approaches to process the formation and splitting of inflecti…

Morphological Analysis

Computational Algorithms Based on the Paninian System to Process Euphonic Conjunctions for Word Searches

2014-09-15 · S. V. Kasmir Raja, V. Rajitha, Meenakshi Lakshmanan

Searching for words in Sanskrit E-text is a problem that is accompanied by complexities introduced by features of Sanskrit such as euphonic conjunctions or sandhis. A word could occur in an E-text in a transformed form o…