paper-with-me

Papers

Pairing Orthographically Variant Literary Words to Standard Equivalents Using Neural Edit Distance Models

2024-01-26 · Craig Messner, Tom Lippincott

We present a novel corpus consisting of orthographically variant words found in works of 19th century U.S. literature annotated with their corresponding "standard" word pair. We train a set of neural edit distance models to pair these variants with their standard forms, and compare the performance of these models to the performance of a set of neural edit distance models trained on a corpus of orthographic errors made by L2 English learners. Finally, we analyze the relative performance of these models in the light of different negative training sample generation strategies, and offer concluding remarks on the unique challenge literary orthographic variation poses to string pairing methodologies.

📄 PDF Abstract BibTeX arXiv:2401.15068

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Cognate-aware morphological segmentation for multilingual neural translation

2018-08-31 · WS 2018 10 · Stig-Arne Grönroos, Sami Virpioja, Mikko Kurimo

This article describes the Aalto University entry to the WMT18 News Translation Shared Task. We participate in the multilingual subtrack with a system trained under the constrained condition to translate from English to …

Translation

Variation of word frequencies in Russian literary texts

2015-03-01 · Vladislav Kargin

We study the variation of word frequencies in Russian literary texts. Our findings indicate that the standard deviation of a word's frequency across texts depends on its average frequency according to a power law with ex…

A Data-Oriented Model of Literary Language

2017-01-12 · EACL 2017 4 · Andreas van Cranenburgh, Rens Bod

We consider the task of predicting how literary a text is, with a gold standard from human ratings. Aside from a standard bigram baseline, we apply rich syntactic tree fragments, mined from the training set, and a series…

model

Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages

2026-03-01 · Kaushal Santosh Bhogale, Tahir Javed, Greeshma Susan John, Dhruv Rathi 외 arxiv

Evaluating ASR systems for Indian languages is challenging due to spelling variations, suffix splitting flexibility, and non-standard spellings in code-mixed words. Traditional Word Error Rate (WER) often presents a blea…

Speech Recognition

lakṣyārtha (Indicated Meaning) of Śabdavyāpāra (Function of a Word) framework from kāvyaśāstra (The Science of Literary Studies) in Samskṛtam : Its application to Literary Machine Translation and other NLP tasks

2021-12-01 · ICON 2021 12 · Sripathi Sripada, Anupama Ryali, Raghuram Sheshadri

A key challenge in Literary Machine Translation is that the meaning of a sentence can be different from the sum of meanings of all the words it possesses. This poses the problem of requiring large amounts of consistently…

Machine TranslationSentenceTranslation