paper-with-me

홈 › Papers

Factorization of Fact-Checks for Low Resource Indian Languages

2021-02-23 · Shivangi Singhal, Rajiv Ratn Shah, Ponnurangam Kumaraguru

The advancement in technology and accessibility of internet to each individual is revolutionizing the real time information. The liberty to express your thoughts without passing through any credibility check is leading to dissemination of fake content in the ecosystem. It can have disastrous effects on both individuals and society as a whole. The amplification of fake news is becoming rampant in India too. Debunked information often gets republished with a replacement description, claiming it to depict some different incidence. To curb such fabricated stories, it is necessary to investigate such deduplicates and false claims made in public. The majority of studies on automatic fact-checking and fake news detection is restricted to English only. But for a country like India where only 10% of the literate population speak English, role of regional languages in spreading falsity cannot be undermined. In this paper, we introduce FactDRIL: the first large scale multilingual Fact-checking Dataset for Regional Indian Languages. We collect an exhaustive dataset across 7 months covering 11 low-resource languages. Our propose dataset consists of 9,058 samples belonging to English, 5,155 samples to Hindi and remaining 8,222 samples are distributed across various regional languages, i.e. Bangla, Marathi, Malayalam, Telugu, Tamil, Oriya, Assamese, Punjabi, Urdu, Sinhala and Burmese. We also present the detailed characterization of three M's (multi-lingual, multi-media, multi-domain) in the FactDRIL accompanied with the complete list of other varied attributes making it a unique dataset to study. Lastly, we present some potential use cases of the dataset. We expect this dataset will be a valuable resource and serve as a starting point to fight proliferation of fake news in low resource languages.

📄 PDF Abstract BibTeX arXiv:2102.11276

Code (0)

등록된 구현이 없습니다.

Tasks

Fact CheckingFake News Detection

Similar Papers 제목 키워드 기반

Indian Regional Movie Dataset for Recommender Systems

2018-01-07 · Prerna Agarwal, Richa Verma, Angshul Majumdar

Indian regional movie dataset is the first database of regional Indian movies, users and their ratings. It consists of movies belonging to 18 different Indian regional languages and metadata of users with varying demogra…

Collaborative Filteringcompressed sensingDiversityMatrix Completion+1

Massively Multilingual Language Models for Cross Lingual Fact Extraction from Low Resource Indian Languages

2023-02-09 · Bhavyajeet Singh, Pavan Kandru, Anubhav Sharma, Vasudeva Varma

Massive knowledge graphs like Wikidata attempt to capture world knowledge about multiple entities. Recent approaches concentrate on automatically enriching these KGs from text. However a lot of information present in the…

FormKnowledge GraphsWorld Knowledge

IndiText Boost: Text Augmentation for Low Resource India Languages

2024-01-23 · Onkar Litake, Niraj Yagnik, Shreyas Labhsetwar

Text Augmentation is an important task for low-resource languages. It helps deal with the problem of data scarcity. A data augmentation strategy is used to deal with the problem of data scarcity. Through the years, much …

Data AugmentationMulti Class Text ClassificationText Augmentationtext-classification+2

Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages

2025-05-27 · Praveen Srinivasa Varadhan, Srija Anand, Soma Siddhartha, Mitesh M. Khapra

What happens when an English Fairytaler is fine-tuned on Indian languages? We evaluate how the English F5-TTS model adapts to 11 Indian languages, measuring polyglot fluency, voice-cloning, style-cloning, and code-mixing…

Synthetic Data GenerationVoice Cloning

Language Resources and Technologies for Non-Scheduled and Endangered Indian Languages

2022-04-06 · Ritesh Kumar, Bornini Lahiri

In the present paper, we will present a survey of the language resources and technologies available for the non-scheduled and endangered languages of India. While there have been different estimates from different source…