paper-with-me

홈 › Papers

AI-Generated Text Detection in Low-Resource Languages: A Case Study on Urdu

2025-10-18 · Muhammad Ammar, Hadiya Murad Hadi, Usman Majeed Butt arxiv

Large Language Models (LLMs) are now capable of generating text that closely resembles human writing, making them powerful tools for content creation, but this growing ability has also made it harder to tell whether a piece of text was written by a human or by a machine. This challenge becomes even more serious for languages like Urdu, where there are very few tools available to detect AI-generated text. To address this gap, we propose a novel AI-generated text detection framework tailored for the Urdu language. A balanced dataset comprising 1,800 humans authored, and 1,800 AI generated texts, sourced from models such as Gemini, GPT-4o-mini, and Kimi AI was developed. Detailed linguistic and statistical analysis was conducted, focusing on features such as character and word counts, vocabulary richness (Type Token Ratio), and N-gram patterns, with significance evaluated through t-tests and MannWhitney U tests. Three state-of-the-art multilingual transformer models such as mdeberta-v3-base, distilbert-base-multilingualcased, and xlm-roberta-base were fine-tuned on this dataset. The mDeBERTa-v3-base achieved the highest performance, with an F1-score 91.29 and accuracy of 91.26% on the test set. This research advances efforts in contesting misinformation and academic misconduct in Urdu-speaking communities and contributes to the broader development of NLP tools for low resource languages.

📄 PDF Abstract BibTeX arXiv:2510.16573

Code (0)

등록된 구현이 없습니다.

Tasks

Text Detection

Similar Papers 제목 키워드 기반

BLUFF: Benchmarking the Detection of False and Synthetic Content across 58 Low-Resource Languages

2026-02-28 · Jason Lucas, Matt Murtagh-White, Adaku Uchendu, Ali Al-Lawati 외 arxiv

Multilingual falsehoods threaten information integrity worldwide, yet detection benchmarks remain confined to English or a few high-resource languages, leaving low-resource linguistic communities without robust defense t…

Can LLMs Faithfully Explain Themselves in Low-Resource Languages? A Case Study on Emotion Detection in Persian

2025-11-24 · Mobina Mehrazar, Mohammad Amin Yousefi, Parisa Abolfath Beygi, Behnam Bahrak arxiv

Large language models (LLMs) are increasingly used to generate self-explanations alongside their predictions, a practice that raises concerns about the faithfulness of these explanations, especially in low-resource langu…

Emotion Classification

Classification of Human- and AI-Generated Texts for English, French, German, and Spanish

2023-12-08 · Kristina Schaaff, Tim Schlippe, Lorenz Mindner

In this paper we analyze features to classify human- and AI-generated text for English, French, German and Spanish and compare them across languages. We investigate two scenarios: (1) The detection of text generated by A…

Findings of the Shared Task on Offensive Language Identification in Tamil, Malayalam, and Kannada

2021-04-01 · EACL (DravidianLangTech) 2021 4 · Bharathi Raja Chakravarthi, Ruba Priyadharshini, Navya Jose, Anand Kumar M 외

Detecting offensive language in social media in local languages is critical for moderating user-generated content. Thus, the field of offensive language identification in under-resourced Tamil, Malayalam and Kannada lang…

BenchmarkingLanguage Identification

SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia

2026-03-17 · Ri Chi Ng, Aditi Kumaresan, Yujia Hu, Roy Ka-Wei Lee arxiv

Hate speech detection relies heavily on linguistic resources, which are primarily available in high-resource languages such as English and Chinese, creating barriers for researchers and platforms developing tools for low…

Hate Speech Detection