paper-with-me

Papers

Romanian Diacritics Restoration Using Recurrent Neural Networks

2020-09-06 · Stefan Ruseti, Teodor-Mihai Cotet, Mihai Dascalu

Diacritics restoration is a mandatory step for adequately processing Romanian texts, and not a trivial one, as you generally need context in order to properly restore a character. Most previous methods which were experimented for Romanian restoration of diacritics do not use neural networks. Among those that do, there are no solutions specifically optimized for this particular language (i.e., they were generally designed to work on many different languages). Therefore we propose a novel neural architecture based on recurrent neural networks that can attend information at different levels of abstractions in order to restore diacritics.

📄 PDF Abstract BibTeX arXiv:2009.02743

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluating Large Language Models for Diacritic Restoration in Romanian Texts: A Comparative Study

2025-11-17 · Mihai Nadas, Laura Diosan arxiv

Automatic diacritic restoration is crucial for text processing in languages with rich diacritical marks, such as Romanian. This study evaluates the performance of several large language models (LLMs) in restoring diacrit…

RoBERT -- A Romanian BERT Model

2020-12-01 · COLING 2020 8 · Mihai Masala, Stefan Ruseti, Mihai Dascalu

Deep pre-trained language models tend to become ubiquitous in the field of Natural Language Processing (NLP). These models learn contextualized representations by using a huge amount of unlabeled text data and obtain sta…

modelSentiment AnalysisTransfer Learning

Diacritics Restoration using BERT with Analysis on Czech language

2021-05-24 · Jakub Náplava, Milan Straka, Jana Straková

We propose a new architecture for diacritics restoration based on contextualized embeddings, namely BERT, and we evaluate it on 12 languages with diacritics. Furthermore, we conduct a detailed error analysis on Czech, a …

Croatian Text DiacritizationCzech Text DiacritizationFrench Text DiacritizationHungarian Text Diacritization+8

Lexical Disambiguation of Igbo using Diacritic Restoration

2017-04-01 · WS 2017 4 · Ignatius Ezeani, Mark Hepple, Ikechukwu Onyenwe

Properly written texts in Igbo, a low-resource African language, are rich in both orthographic and tonal diacritics. Diacritics are essential in capturing the distinctions in pronunciation and meaning of words, as well a…

BIG-bench Machine LearningGeneral Classification

Dialectal and Low-Resource Machine Translation for Aromanian

2024-10-23 · Alexandru-Iulius Jerpelea, Alina Rădoi, Sergiu Nisioi

This paper presents the process of building a neural machine translation system with support for English, Romanian, and Aromanian - an endangered Eastern Romance language. The primary contribution of this research is two…

Machine TranslationSentenceSentence EmbeddingSentence-Embedding+1