paper-with-me

Papers

Creation of comparable corpora for English-Urdu, Arabic, Persian

2016-05-01 · LREC 2016 5 · Murad Abouammoh, Kashif Shah, Ahmet Aker

Statistical Machine Translation (SMT) relies on the availability of rich parallel corpora. However, in the case of under-resourced languages or some specific domains, parallel corpora are not readily available. This leads to under-performing machine translation systems in those sparse data settings. To overcome the low availability of parallel resources the machine translation community has recognized the potential of using comparable resources as training data. However, most efforts have been related to European languages and less in middle-east languages. In this study, we report comparable corpora created from news articles for the pair English ―{Arabic, Persian, Urdu} languages. The data has been collected over a period of a year, entails Arabic, Persian and Urdu languages. Furthermore using the English as a pivot language, comparable corpora that involve more than one language can be created, e.g. English- Arabic - Persian, English - Arabic - Urdu, English ― Urdu - Persian, etc. Upon request the data can be provided for research purposes.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesMachine TranslationTranslation

Similar Papers 제목 키워드 기반

Learning Trilingual Dictionaries for Urdu -- Roman Urdu -- English

2019-08-01 · WS 2019 8 · Moiz Rauf, Sebastian Pad{\'o}

In this paper, we present an effort to generate a joint Urdu, Roman Urdu and English trilingual lexicon using automated methods. We make a case for using statistical machine translation approaches and parallel corpora fo…

Machine TranslationTranslationWord Alignment

Data Augmentation using Machine Translation for Fake News Detection in the Urdu Language

2020-05-01 · LREC 2020 5 · Maaz Amjad, Grigori Sidorov, Alisa Zhila

The task of fake news detection is to distinguish legitimate news articles that describe real facts from those which convey deceiving and fictitious information. As the fake news phenomenon is omnipresent across all lang…

ArticlesData AugmentationFake News DetectionMachine Translation+1

PALO: A Polyglot Large Multimodal Model for 5B People

2024-02-22 · Muhammad Maaz, Hanoona Rasheed, Abdelrahman Shaker, Salman Khan 외

In pursuit of more inclusive Vision-Language Models (VLMs), this study introduces a Large Multilingual Multimodal Model called PALO. PALO offers visual reasoning capabilities in 10 major languages, including English, Chi…

Language ModelingLanguage ModellingLarge Language ModelVisual Reasoning

Rule Based Stemmer in Urdu

2013-10-02 · Vaishali Gupta, Nisheeth Joshi, Iti Mathur

Urdu is a combination of several languages like Arabic, Hindi, English, Turkish, Sanskrit etc. It has a complex and rich morphology. This is the reason why not much work has been done in Urdu language processing. Stemmin…

Information RetrievalRetrieval

Sentiment Classification of Customer Reviews about Automobiles in Roman Urdu

2018-12-30 · Moin Khan, Kamran Malik

Text mining is a broad field having sentiment mining as its important constituent in which we try to deduce the behavior of people towards a specific item, merchandise, politics, sports, social media comments, review sit…

AttributeClassificationGeneral ClassificationSentiment Analysis+3