Data Augmentation using Machine Translation for Fake News Detection in the Urdu Language
The task of fake news detection is to distinguish legitimate news articles that describe real facts from those which convey deceiving and fictitious information. As the fake news phenomenon is omnipresent across all languages, it is crucial to be able to efficiently solve this problem for languages other than English. A common approach to this task is supervised classification using features of various complexity. Yet supervised machine learning requires substantial amount of annotated data. For English and a small number of other languages, annotated data availability is much higher, whereas for the vast majority of languages, it is almost scarce. We investigate whether machine translation at its present state could be successfully used as an automated technique for annotated corpora creation and augmentation for fake news detection focusing on the English-Urdu language pair. We train a fake news classifier for Urdu on (1) the manually annotated dataset originally in Urdu and (2) the machine-translated version of an existing annotated fake news dataset originally in English. We show that at the present state of machine translation quality for the English-Urdu language pair, the fully automated data augmentation through machine translation did not provide improvement for fake news detection in Urdu.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesData AugmentationFake News DetectionMachine TranslationTranslationSimilar Papers 제목 키워드 기반
Tackling Fake News in Bengali: Unraveling the Impact of Summarization vs. Augmentation on Pre-trained Language Models
With the rise of social media and online news sources, fake news has become a significant issue globally. However, the detection of fake news in low resource languages like Bengali has received limited attention in resea…
ArticlesFake News DetectionAddressing Data Scarcity in Bangla Fake News Detection: An LLM-Based Dataset Augmentation Approach
The growing spread of misinformation in digital media highlights the need for reliable fake news detection systems, yet progress in under-resourced languages such as Bangla is limited by small and imbalanced datasets. Th…
Fake News DetectionNews ClassificationAdversarial Style Augmentation via Large Language Model for Robust Fake News Detection
The spread of fake news harms individuals and presents a critical social challenge that must be addressed. Although numerous algorithmic and insightful features have been developed to detect fake news, many of these feat…
Fake News DetectionLanguage ModelingLanguage ModellingLarge Language ModelThe use of Data Augmentation as a technique for improving neural network accuracy in detecting fake news about COVID-19
This paper aims to present how the application of Natural Language Processing (NLP) and data augmentation techniques can improve the performance of a neural network for better detection of fake news in the Portuguese lan…
Data AugmentationNews ClassificationAdapting Fake News Detection to the Era of Large Language Models
In the age of large language models (LLMs) and the widespread adoption of AI-driven content creation, the landscape of information dissemination has witnessed a paradigm shift. With the proliferation of both human-writte…
ArticlesFake News Detection