paper-with-me

Papers

Learning Trilingual Dictionaries for Urdu -- Roman Urdu -- English

2019-08-01 · WS 2019 8 · Moiz Rauf, Sebastian Pad{\'o}

In this paper, we present an effort to generate a joint Urdu, Roman Urdu and English trilingual lexicon using automated methods. We make a case for using statistical machine translation approaches and parallel corpora for dictionary creation. To this purpose, we use word alignment tools on the corpus and evaluate translations using human evaluators. Despite different writing script and considerable noise in the corpus our results show promise with over 85{\%} accuracy of Roman Urdu{--}Urdu and 45{\%} English{--}Urdu pairs.

📄 PDF Abstract BibTeX

Code (1)

MoizRauf/Urdu--Roman-Urdu--English--Dictionary 공식 구현

Tasks

Machine TranslationTranslationWord Alignment

Similar Papers 제목 키워드 기반

RUBERT: A Bilingual Roman Urdu BERT Using Cross Lingual Transfer Learning

2021-02-22 · Usama Khalid, Mirza Omer Beg, Muhammad Umair Arshad

In recent studies, it has been shown that Multilingual language models underperform their monolingual counterparts. It is also a well-known fact that training and maintaining monolingual models for each language is a cos…

Cross-Lingual TransferTransfer Learning

Fine-Tuning Large Language Models with QLoRA for Offensive Language Detection in Roman Urdu-English Code-Mixed Text

2025-10-04 · Nisar Hussain, Amna Qasim, Gull Mehak, Muhammad Zain 외 arxiv

The use of derogatory terms in languages that employ code mixing, such as Roman Urdu, presents challenges for Natural Language Processing systems due to unstated grammar, inconsistent spelling, and a scarcity of labeled …

ERUPD -- English to Roman Urdu Parallel Dataset

2024-12-23 · Mohammed Furqan, Raahid Bin Khaja, Rayyan Habeeb

Bridging linguistic gaps fosters global growth and cultural exchange. This study addresses the challenges of Roman Urdu -- a Latin-script adaptation of Urdu widely used in digital communication -- by creating a novel par…

Machine TranslationPrompt EngineeringSentenceSentiment Analysis

Evaluating Large Language Models on Urdu Idiom Translation

2025-10-20 · Muhammad Farmal Khan, Mousumi Akter arxiv

Idiomatic translation remains a significant challenge in machine translation, especially for low resource languages such as Urdu, and has received limited prior attention. To advance research in this area, we introduce t…

Machine TranslationPrompt Engineering

A Clustering Framework for Lexical Normalization of Roman Urdu

2020-03-31 · Abdul Rafae Khan, Asim Karim, Hassan Sajjad, Faisal Kamiran 외

Roman Urdu is an informal form of the Urdu language written in Roman script, which is widely used in South Asia for online textual content. It lacks standard spelling and hence poses several normalization challenges duri…

ClusteringLexical Normalization