paper-with-me

Papers

Are LLMs Ready to Replace Bangla Annotators?

2026-02-18 · Md. Najib Hasan, Touseef Hasan, Souvika Sarkar arxiv

Large Language Models (LLMs) are increasingly used as automated annotators to scale dataset creation, yet their reliability as unbiased annotators--especially for low-resource and identity-sensitive settings--remains poorly understood. In this work, we study the behavior of LLMs as zero-shot annotators for Bangla hate speech, a task where even human agreement is challenging, and annotator bias can have serious downstream consequences. We conduct a systematic benchmark of 17 LLMs using a unified evaluation framework. Our analysis uncovers annotator bias and substantial instability in model judgments. Surprisingly, increased model scale does not guarantee improved annotation quality--smaller, more task-aligned models frequently exhibit more consistent behavior than their larger counterparts. These results highlight important limitations of current LLMs for sensitive annotation tasks in low-resource languages and underscore the need for careful evaluation before deployment.

📄 PDF Abstract BibTeX arXiv:2602.16241

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BAN-Cap: A Multi-Purpose English-Bangla Image Descriptions Dataset

2022-05-28 · LREC 2022 6 · Mohammad Faiyaz Khan, S. M. Sadiq-Ur-Rahman Shifath, Md Saiful Islam

As computers have become efficient at understanding visual information and transforming it into a written representation, research interest in tasks like automatic image captioning has seen a significant leap over the la…

Image CaptioningMachine TranslationText Augmentation

Can LLMs Follow the Pulse of a Crisis? Evaluating Crisis Sentiment in Bangladesh's July Uprising

2026-09-15 · Md. Samiul Alim, Mahir Shahriar Tamim, Tanvir Ahmed Khan, Sharjil Khan 외 arxiv

Crisis sentiment analysis is especially challenging for low-resource languages such as Bangla, where language, context, and public reaction shift rapidly. We introduce UNRESTSENT200K, a Bangla crisis sentiment dataset wi…

Sentiment Analysis

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

2025-01-19 · Nitay Calderon, Roi Reichart, Rotem Dror

The "LLM-as-an-annotator" and "LLM-as-a-judge" paradigms employ Large Language Models (LLMs) as annotators, judges, and evaluators in tasks traditionally performed by humans. LLM annotations are widely used, not only in …

LLMs for Qualitative Data Analysis Fail on Security-specificComments in Human Experiments

2026-04-12 · Maria Camporese, Fabio Massacci, Yuanjun Gong arxiv

[Background:] Thematic analysis of free-text justifications in human experiments provides significant qualitative insights. Yet, it is costly because reliable annotations require multiple domain experts. Large language m…

TigerCoder: A Novel Suite of LLMs for Code Generation in Bangla

2025-09-11 · Nishat Raihan, Antonios Anastasopoulos, Marcos Zampieri arxiv

Despite being the 5th most spoken language, Bangla remains underrepresented in Large Language Models (LLMs), particularly for code generation. This primarily stems from the scarcity of high-quality data to pre-train and/…

Domain AdaptationCode Generation