paper-with-me

홈 › Papers

Compound Type Identification in Sanskrit: What Roles do the Corpus and Grammar Play?

2016-12-01 · WS 2016 12 · Amrith Krishna, Pavankumar Satuluri, Shubham Sharma, Apurv Kumar, Pawan Goyal

We propose a classification framework for semantic type identification of compounds in Sanskrit. We broadly classify the compounds into four different classes namely, \textit{Avyay{\=\i}bh{\=a}va}, \textit{Tatpuruṣa}, \textit{Bahuvr{\=\i}hi} and \textit{Dvandva}. Our classification is based on the traditional classification system followed by the ancient grammar treatise \textit{Adṣṭ{\=a}dhy{\=a}y{\=\i}}, proposed by P{\=a}ṇini 25 centuries back. We construct an elaborate features space for our system by combining conditional rules from the grammar \textit{Adṣṭ{\=a}dhy{\=a}y{\=\i}}, semantic relations between the compound components from a lexical database \textit{Amarakoṣa} and linguistic structures from the data using Adaptor Grammars. Our in-depth analysis of the feature space highlight inadequacy of \textit{Adṣṭ{\=a}dhy{\=a}y{\=\i}}, a generative grammar, in classifying the data samples. Our experimental results validate the effectiveness of using lexical databases as suggested by Amba Kulkarni and Anil Kumar, and put forward a new research direction by introducing linguistic patterns obtained from Adaptor grammars for effective identification of compound type. We utilise an ensemble based approach, specifically designed for handling skewed datasets and we {\%}and Experimenting with various classification methods, we achieve an overall accuracy of 0.77 using random forest classifiers.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral ClassificationMachine TranslationQuestion Answering

Similar Papers 제목 키워드 기반

DepNeCTI: Dependency-based Nested Compound Type Identification for Sanskrit

2023-10-14 · Jivnesh Sandhan, Yaswanth Narsupalli, Sreevatsa Muppirala, Sriram Krishnan 외

Multi-component compounding is a prevalent phenomenon in Sanskrit, and understanding the implicit structure of a compound's components is crucial for deciphering its meaning. Earlier approaches in Sanskrit have focused o…

Constituency Parsingnamed-entity-recognitionNamed Entity RecognitionNested Named Entity Recognition

Revisiting the Role of Feature Engineering for Compound Type Identification in Sanskrit

2019-10-01 · WS 2019 10 · S, Jivnesh han, Amrith Krishna, Pawan Goyal 외
Feature Engineering

Linguistically-Informed Neural Architectures for Lexical, Syntactic and Semantic Tasks in Sanskrit

2023-08-17 · Jivnesh Sandhan

The primary focus of this thesis is to make Sanskrit manuscripts more accessible to the end-users through natural language technologies. The morphological richness, compounding, free word orderliness, and low-resource na…

Dependency ParsingMachine TranslationQuestion Answering

A Novel Multi-Task Learning Approach for Context-Sensitive Compound Type Identification in Sanskrit

2022-08-22 · COLING 2022 10 · Jivnesh Sandhan, Ashish Gupta, Hrishikesh Terdalkar, Tushar Sandhan 외

The phenomenon of compounding is ubiquitous in Sanskrit. It serves for achieving brevity in expressing thoughts, while simultaneously enriching the lexical and structural formation of the language. In this work, we focus…

Dependency ParsingMorphological TaggingMulti-Task Learning

SanskritShala: A Neural Sanskrit NLP Toolkit with Web-Based Interface for Pedagogical and Annotation Purposes

2023-02-19 · Jivnesh Sandhan, Anshul Agarwal, Laxmidhar Behera, Tushar Sandhan 외

We present a neural Sanskrit Natural Language Processing (NLP) toolkit named SanskritShala (a school of Sanskrit) to facilitate computational linguistic analyses for several tasks such as word segmentation, morphological…

Dependency ParsingMorphological TaggingWord EmbeddingsWord Similarity