paper-with-me

홈 › Papers

A Hybrid Algorithm for Matching Arabic Names

2013-09-22 · T. El-Shishtawy

In this paper, a new hybrid algorithm which combines both of token-based and character-based approaches is presented. The basic Levenshtein approach has been extended to token-based distance metric. The distance metric is enhanced to set the proper granularity level behavior of the algorithm. It smoothly maps a threshold of misspellings differences at the character level, and the importance of token level errors in terms of token's position and frequency. Using a large Arabic dataset, the experimental results show that the proposed algorithm overcomes successfully many types of errors such as: typographical errors, omission or insertion of middle name components, omission of non-significant popular name components, and different writing styles character variations. When compared the results with other classical algorithms, using the same dataset, the proposed algorithm was found to increase the minimum success level of best tested algorithms, while achieving higher upper limits .

📄 PDF Abstract BibTeX arXiv:1309.5657

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Cross-Language Personal Name Mapping

2014-05-24 · Ahmed H. Yousef

Name matching between multiple natural languages is an important step in cross-enterprise integration applications and data mining. It is difficult to decide whether or not two syntactic values (names) from two heterogen…

Rule-and Dictionary-based Solution for Variations in Written Arabic Names in Social Networks, Big Data, Accounting Systems and Large Databases

2015-02-18 · Ahmad B. A. Hassanat, Ghada Awad Altarawneh

This paper investigates the problem that some Arabic names can be written in multiple ways. When someone searches for only one form of a name, neither exact nor approximate matching is appropriate for returning the multi…

A hybrid approach to Vietnamese word segmentation

2016-12-29 · Tuan-Phong Nguyen, Anh-Cuong Le

Word segmentation is the very first task for Vietnamese language processing. Word-segmented text is the input of almost other NLP tasks. This task faces some challenges due to specific characteristics of the language. As…

SegmentationSentenceVietnamese Word Segmentation

Using Transliteration of Proper Names from Arabic to Latin Script to Improve English-Arabic Word Alignment

2013-10-01 · IJCNLP 2013 10 · Nasredine Semmar, Houda Saadane
Information RetrievalMachine TranslationTransliterationWord Alignment

Proper Name Diacritization for Arabic Wikipedia: A Benchmark Dataset

2025-05-05 · Rawan Bondok, Mayar Nassar, Salam Khalifa, Kurt Micallef 외

Proper names in Arabic Wikipedia are frequently undiacritized, creating ambiguity in pronunciation and interpretation, especially for transliterated named entities of foreign origin. While transliteration and diacritizat…

Transliteration