paper-with-me

홈 › Papers

Pragya: An AI-Based Semantic Recommendation System for Sanskrit Subhasitas

2026-01-10 · Tanisha Raorane, Prasenjit Kole arxiv

Sanskrit Subhasitas encapsulate centuries of cultural and philosophical wisdom, yet remain underutilized in the digital age due to linguistic and contextual barriers. In this work, we present Pragya, a retrieval-augmented generation (RAG) framework for semantic recommendation of Subhasitas. We curate a dataset of 200 verses annotated with thematic tags such as motivation, friendship, and compassion. Using sentence embeddings (IndicBERT), the system retrieves top-k verses relevant to user queries. The retrieved results are then passed to a generative model (Mistral LLM) to produce transliterations, translations, and contextual explanations. Experimental evaluation demonstrates that semantic retrieval significantly outperforms keyword matching in precision and relevance, while user studies highlight improved accessibility through generated summaries. To our knowledge, this is the first attempt at integrating retrieval and generation for Sanskrit Subhasitas, bridging cultural heritage with modern applied AI.

📄 PDF Abstract BibTeX arXiv:2601.06607

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Retrieval

Similar Papers 제목 키워드 기반

Abstractive Text Summarization for Sanskrit Prose: A Study of Methods and Approaches

2020-05-01 · LREC 2020 5 · Shagun Sinha, Girish Jha

The authors present a work-in-progress in the field of Abstractive Text Summarization (ATS) for Sanskrit Prose {--} a first attempt at ATS for Sanskrit (SATS). We will evaluate recent approaches and methods used for ATS …

Abstractive Text SummarizationExtractive Text SummarizationInformation RetrievalRetrieval+1

Linguistically-Informed Neural Architectures for Lexical, Syntactic and Semantic Tasks in Sanskrit

2023-08-17 · Jivnesh Sandhan

The primary focus of this thesis is to make Sanskrit manuscripts more accessible to the end-users through natural language technologies. The morphological richness, compounding, free word orderliness, and low-resource na…

Dependency ParsingMachine TranslationQuestion Answering

Pragyaan: Designing and Curating High-Quality Cultural Post-Training Datasets for Indian Languages

2025-10-08 · Neel Prabhanjan Rachamalla, Aravind Konakalla, Gautam Rajeev, Ashish Kulkarni 외 arxiv

The effectiveness of Large Language Models (LLMs) depends heavily on the availability of high-quality post-training data, particularly instruction-tuning and preference-based examples. Existing open-source datasets, howe…

Sanskrit Word Segmentation Using Character-level Recurrent and Convolutional Neural Networks

2018-10-01 · EMNLP 2018 10 · Oliver Hellwig, Sebastian Nehrdich

The paper introduces end-to-end neural network models that tokenize Sanskrit by jointly splitting compounds and resolving phonetic merges (Sandhi). Tokenization of Sanskrit depends on local phonetic and distant semantic …

Feature Engineering

An evaluation of Google Translate for Sanskrit to English translation via sentiment and semantic analysis

2023-02-28 · Akshat Shukla, Chaarvi Bansal, Sushrut Badhe, Mukul Ranjan 외

Google Translate has been prominent for language translation; however, limited work has been done in evaluating the quality of translation when compared to human experts. Sanskrit one of the oldest written languages in t…

Translation