paper-with-me

홈 › Papers

Comprehensive Benchmark Datasets for Amharic Scene Text Detection and Recognition

2022-03-23 · Wondimu Dikubab, Dingkang Liang, Minghui Liao, Xiang Bai

Ethiopic/Amharic script is one of the oldest African writing systems, which serves at least 23 languages (e.g., Amharic, Tigrinya) in East Africa for more than 120 million people. The Amharic writing system, Abugida, has 282 syllables, 15 punctuation marks, and 20 numerals. The Amharic syllabic matrix is derived from 34 base graphemes/consonants by adding up to 12 appropriate diacritics or vocalic markers to the characters. The syllables with a common consonant or vocalic markers are likely to be visually similar and challenge text recognition tasks. In this work, we presented the first comprehensive public datasets named HUST-ART, HUST-AST, ABE, and Tana for Amharic script detection and recognition in the natural scene. We have also conducted extensive experiments to evaluate the performance of the state of art methods in detecting and recognizing Amharic scene text on our datasets. The evaluation results demonstrate the robustness of our datasets for benchmarking and its potential of promoting the development of robust Amharic script detection and recognition algorithms. Consequently, the outcome will benefit people in East Africa, including diplomats from several countries and international communities.

📄 PDF Abstract BibTeX arXiv:2203.12165

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingScene Text DetectionText Detection

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

AmQA: Amharic Question Answering Dataset

2023-03-06 · Tilahun Abedissa, Ricardo Usbeck, Yaregal Assabie

Question Answering (QA) returns concise answers or answer lists from natural language text given a context document. Many resources go into curating QA datasets to advance robust models' development. There is a surge of …

ArticlesQuestion AnsweringReading Comprehension

AmharicIR+Instr: A Two-Dataset Resource for Neural Retrieval and Instruction Tuning

2026-02-10 · Tilahun Yeshambel, Moncef Garouani, Josiane Mothe arxiv

Neural retrieval and GPT-style generative models rely on large, high-quality supervised data, which is still scarce for low-resource languages such as Amharic. We release an Amharic data resource consisting of two datase…

Text Generation

Optimized Text Embedding Models and Benchmarks for Amharic Passage Retrieval

2025-05-25 · Kidist Amde Mekonnen, Yosef Worku Alemneh, Maarten de Rijke

Neural retrieval methods using transformer-based pre-trained language models have advanced multilingual and cross-lingual retrieval. However, their effectiveness for low-resource, morphologically rich languages such as A…

Passage RetrievalRetrieval

AmaSQuAD: A Benchmark for Amharic Extractive Question Answering

2025-02-04 · Nebiyou Daniel Hailemariam, Blessed Guda, Tsegazeab Tefferi

This research presents a novel framework for translating extractive question-answering datasets into low-resource languages, as demonstrated by the creation of the AmaSQuAD dataset, a translation of SQuAD 2.0 into Amhari…

Extractive Question-AnsweringQuestion AnsweringXLM-R

Semantically Corrected Amharic Automatic Speech Recognition

2024-04-20 · Samuael Adnew, Paul Pu Liang

Automatic Speech Recognition (ASR) can play a crucial role in enhancing the accessibility of spoken languages worldwide. In this paper, we build a set of ASR tools for Amharic, a language spoken by more than 50 million p…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderSentence+2