paper-with-me

홈 › Papers

BTPD: A Multilingual Hand-curated Dataset of Bengali Transnational Political Discourse Across Online Communities

2025-06-07 · Dipto Das, Syed Ishtiaque Ahmed, Shion Guha

Understanding political discourse in online spaces is crucial for analyzing public opinion and ideological polarization. While social computing and computational linguistics have explored such discussions in English, such research efforts are significantly limited in major yet under-resourced languages like Bengali due to the unavailability of datasets. In this paper, we present a multilingual dataset of Bengali transnational political discourse (BTPD) collected from three online platforms, each representing distinct community structures and interaction dynamics. Besides describing how we hand-curated the dataset through community-informed keyword-based retrieval, this paper also provides a general overview of its topics and multilingual content.

📄 PDF Abstract BibTeX arXiv:2506.06813

Code (0)

등록된 구현이 없습니다.

Tasks

Retrieval

Similar Papers 제목 키워드 기반

BNLI: A Linguistically-Refined Bengali Dataset for Natural Language Inference

2025-11-11 · Farah Binta Haque, Md Yasin, Shishir Saha, Md Shoaib Akhter Rafi 외 arxiv

Despite the growing progress in Natural Language Inference (NLI) research, resources for the Bengali language remain extremely limited. Existing Bengali NLI datasets exhibit several inconsistencies, including annotation …

Natural Language Inference

From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge

2025-10-22 · Nafis Chowdhury, Moinul Haque, Anika Ahmed, Nazia Tasnim 외 arxiv

Recent progress in NLP research has demonstrated remarkable capabilities of large language models (LLMs) across a wide range of tasks. While recent multilingual benchmarks have advanced cultural evaluation for LLMs, crit…

A Large Multi-Target Dataset of Common Bengali Handwritten Graphemes

2020-10-01 · Samiul Alam, Tahsin Reasat, Asif Shahriyar Sushmit, Sadi Mohammad Siddiquee 외

Latin has historically led the state-of-the-art in handwritten optical character recognition (OCR) research. Adapting existing systems from Latin to alpha-syllabary languages is particularly challenging due to a sharp co…

Multi-Label ClassificationOptical Character RecognitionOptical Character Recognition (OCR)

KrishokBondhu: A Retrieval-Augmented Voice-Based Agricultural Advisory Call Center for Bengali Farmers

2025-10-21 · Mohd Ruhul Ameen, Akif Islam, Farjana Aktar, M. Saifuzzaman Rafat arxiv

In Bangladesh, many farmers still struggle to access timely, expert-level agricultural guidance. This paper presents KrishokBondhu, a voice-enabled, call-centre-integrated advisory platform built on a Retrieval-Augmented…

Semantic Retrieval

Modeling Romanized Hindi and Bengali: Dataset Creation and Multilingual LLM Integration

2025-11-27 · Kanchon Gharami, Quazi Sarwar Muhtaseem, Deepti Gupta, Lavanya Elluri 외 arxiv

The development of robust transliteration techniques to enhance the effectiveness of transforming Romanized scripts into native scripts is crucial for Natural Language Processing tasks, including sentiment analysis, spee…

Information RetrievalSpeech RecognitionSentiment Analysis