paper-with-me

Papers

Hebrew Diacritics Restoration using Visual Representation

2025-10-30 · Yair Elboher, Yuval Pinter arxiv

Diacritics restoration in Hebrew is a fundamental task for ensuring accurate word pronunciation and disambiguating textual meaning. Despite the language's high degree of ambiguity when unvocalized, recent machine learning approaches have significantly advanced performance on this task. In this work, we present DiVRit, a novel system for Hebrew diacritization that frames the task as a zero-shot classification problem. Our approach operates at the word level, selecting the most appropriate diacritization pattern for each undiacritized word from a dynamically generated candidate set, conditioned on the surrounding textual context. A key innovation of DiVRit is its use of a Hebrew Visual Language Model to process diacritized candidates as images, allowing diacritic information to be embedded directly within their vector representations while the surrounding context remains tokenization-based. Through a comprehensive evaluation across various configurations, we demonstrate that the system effectively performs diacritization without relying on complex, explicit linguistic analysis. Notably, in an ``oracle'' setting where the correct diacritized form is guaranteed to be among the provided candidates, DiVRit achieves a high level of accuracy. Furthermore, strategic architectural enhancements and optimized training methodologies yield significant improvements in the system's overall generalization capabilities. These findings highlight the promising potential of visual representations for accurate and automated Hebrew diacritization.

📄 PDF Abstract BibTeX arXiv:2510.26521

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Language Modeling Approach to Diacritic-Free Hebrew TTS

2024-07-16 · Amit Roth, Arnon Turetzky, Yossi Adi

We tackle the task of text-to-speech (TTS) in Hebrew. Traditional Hebrew contains Diacritics, which dictate the way individuals should pronounce given words, however, modern Hebrew rarely uses them. The lack of diacritic…

Language ModelingLanguage Modellingtext-to-speechText to Speech

A Novel Challenge Set for Hebrew Morphological Disambiguation and Diacritics Restoration

2020-10-06 · Findings of the Association for Computational Linguistics 2020 · Avi Shmidman, Joshua Guedalia, Shaltiel Shmidman, Moshe Koppel 외

One of the primary tasks of morphological parsers is the disambiguation of homographs. Particularly difficult are cases of unbalanced ambiguity, where one of the possible analyses is far more frequent than the others. In…

Morphological Disambiguation

Restoring Hebrew Diacritics Without a Dictionary

2021-05-11 · Findings (NAACL) 2022 7 · Elazar Gershuni, Yuval Pinter

We demonstrate that it is feasible to diacritize Hebrew script without any human-curated resources other than plain diacritized text. We present NAKDIMON, a two-layer character level LSTM, that performs on par with much …

Restoring Hebrew Diacritics Without a Dictionary

2021-06-16 · ACL ARR Jun 2021 6 · Anonymous

We demonstrate that it is feasible to diacritize Hebrew script without any human-curated resources other than plain diacritized text. We present NAKDIMON, a two-layer character level LSTM, that performs on par with much …

Diacritics Restoration using BERT with Analysis on Czech language

2021-05-24 · Jakub Náplava, Milan Straka, Jana Straková

We propose a new architecture for diacritics restoration based on contextualized embeddings, namely BERT, and we evaluate it on 12 languages with diacritics. Furthermore, we conduct a detailed error analysis on Czech, a …

Croatian Text DiacritizationCzech Text DiacritizationFrench Text DiacritizationHungarian Text Diacritization+8