paper-with-me

Papers

Stemmers for Tamil Language: Performance Analysis

2013-10-02 · M. Thangarasu, R. Manavalan

Stemming is the process of extracting root word from the given inflection word and also plays significant role in numerous application of Natural Language Processing (NLP). Tamil Language raises several challenges to NLP, since it has rich morphological patterns than other languages. The rule based approach light-stemmer is proposed in this paper, to find stem word for given inflection Tamil word. The performance of proposed approach is compared to a rule based suffix removal stemmer based on correctly and incorrectly predicted. The experimental result clearly show that the proposed approach light stemmer for Tamil language perform better than suffix removal stemmer and also more effective in Information Retrieval System (IRS).

📄 PDF Abstract BibTeX arXiv:1310.0754

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalRetrieval

Similar Papers 제목 키워드 기반

Bag \& Tag'em - A New Dutch Stemmer

2020-05-01 · LREC 2020 5 · Anne Jonker, Corn{\'e} de Ruijt, Jornt de Gruijl

We propose a novel stemming algorithm that is both robust and accurate compared to state-of-the-art solutions, yet addresses several of the problems that current stemmers face in the Dutch language. The main issue is tha…

TAG

A new hybrid stemming algorithm for Persian

2015-07-11 · Adel Rahimi

Stemming has been an influential part in Information retrieval and search engines. There have been tremendous endeavours in making stemmer that are both efficient and accurate. Stemmers can have three method in stemming,…

Information RetrievalRetrieval

Comparing Apples to Apple: The Effects of Stemmers on Topic Models

2016-01-01 · TACL 2016 1 · Alex Schofield, ra, David Mimno

Rule-based stemmers such as the Porter stemmer are frequently used to preprocess English corpora for topic modeling. In this work, we train and evaluate topic models on a variety of corpora using several different stemmi…

Information RetrievalSemantic Textual SimilarityTopic Models

Systematic Analysis of Large Language Models and Transformer-Based Machine Translation for English-Tamil and Tamil-English Across Diverse Datasets

2026-07-27 · Sriharshaa S, Sangeetha Sivanesan, Jaya Nirmala S arxiv

The challenge of Machine Translation for low resource languages such as Tamil is primarily caused by the restricted amount of parallel data for these languages, as well as their substantial amount of domain variation and…

Machine Translation

Zero-shot OCR Accuracy of Low-Resourced Languages: A Comparative Analysis on Sinhala and Tamil

2025-07-24 · Nevidu Jayatilleke, Nisansa de Silva arxiv

Solving the problem of Optical Character Recognition (OCR) on printed text for Latin and its derivative scripts can now be considered settled due to the volumes of research done on English and other High-Resourced Langua…

Document AI