paper-with-me

Papers

IndRegBias: A Dataset for Studying Indian Regional Biases in English and Code-Mixed Social Media Comments

2026-01-10 · Debasmita Panda, Akash Anil, Neelesh Kumar Shukla arxiv

Warning: This paper consists of examples representing regional biases in Indian regions that might be offensive towards a particular region. While social biases corresponding to gender, race, socio-economic conditions, etc., have been extensively studied in the major applications of Natural Language Processing (NLP), biases corresponding to regions have garnered less attention. This is mainly because of (i) difficulty in the extraction of regional bias datasets, (ii) disagreements in annotation due to inherent human biases, and (iii) regional biases being studied in combination with other types of social biases and often being under-represented. This paper focuses on creating a dataset IndRegBias, consisting of regional biases in an Indian context reflected in users' comments on popular social media platforms, namely Reddit and YouTube. We carefully selected 25,000 comments appearing on various threads in Reddit and videos on YouTube discussing trending topics on regional issues in India. Furthermore, we propose a multilevel annotation strategy to annotate the comments describing the severity of regional biased statements. To detect the presence of regional bias and its severity in IndRegBias, we evaluate open-source Large Language Models (LLMs) and Indic Language Models (ILMs) using zero-shot, few-shot, and fine-tuning strategies. We observe that zero-shot and few-shot approaches show lower accuracy in detecting regional biases and severity in the majority of the LLMs and ILMs. However, the fine-tuning approach significantly enhances the performance of the LLM in detecting Indian regional bias along with its severity.

📄 PDF Abstract BibTeX arXiv:2601.06477

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Indian Regional Movie Dataset for Recommender Systems

2018-01-07 · Prerna Agarwal, Richa Verma, Angshul Majumdar

Indian regional movie dataset is the first database of regional Indian movies, users and their ratings. It consists of movies belonging to 18 different Indian regional languages and metadata of users with varying demogra…

Collaborative Filteringcompressed sensingDiversityMatrix Completion+1

Global Voices, Local Biases: Socio-Cultural Prejudices across Languages

2023-10-26 · Anjishnu Mukherjee, Chahat Raj, Ziwei Zhu, Antonios Anastasopoulos

Human biases are ubiquitous but not uniform: disparities exist across linguistic, cultural, and societal borders. As large amounts of recent literature suggest, language models (LMs) trained on human data can reflect and…

On Evaluating and Mitigating Gender Biases in Multilingual Settings

2023-07-04 · Aniket Vashishtha, Kabir Ahuja, Sunayana Sitaram

While understanding and removing gender biases in language models has been a long-standing problem in Natural Language Processing, prior research work has primarily been limited to English. In this work, we investigate s…

IRLCov19: A Large COVID-19 Multilingual Twitter Dataset of Indian Regional Languages

2021-07-26 · Deepak Uniyal, Amit Agarwal

Emerged in Wuhan city of China in December 2019, COVID-19 continues to spread rapidly across the world despite authorities having made available a number of vaccines. While the coronavirus has been around for a significa…

BharatBBQ: A Multilingual Bias Benchmark for Question Answering in the Indian Context

2025-08-09 · Aditya Tomar, Nihar Ranjan Sahoo, Pushpak Bhattacharyya arxiv

Evaluating social biases in language models (LMs) is crucial for ensuring fairness and minimizing the reinforcement of harmful stereotypes in AI systems. Existing benchmarks, such as the Bias Benchmark for Question Answe…

Question Answering