paper-with-me

홈 › Papers

BnPC: A Corpus for Paraphrase Detection in Bangla

2021-12-17 · ACL ARR December 2022 12 · Anonymous

In this paper, we present the first benchmark dataset for paraphrase detection in Bangla language. Despite being the sixth most spoken language in the world, paraphrase identification in the Bangla language is barely explored. Our dataset contains 8,787 human-annotated sentence pairs collected from a total of 23 newspaper outlets' headlines on four categories. We explore different linguistic features and pre-trained language models to benchmark the dataset. We perform a human evaluation experiment to obtain a better understanding of the task's constraints, which reveals intriguing insights. We make our dataset and code publicly available.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Paraphrase IdentificationSentence

Similar Papers 제목 키워드 기반

BanglaParaphrase: A High-Quality Bangla Paraphrase Dataset

2022-10-11 · Ajwad Akil, Najrin Sultana, Abhik Bhattacharjee, Rifat Shahriyar

In this work, we present BanglaParaphrase, a high-quality synthetic Bangla Paraphrase dataset curated by a novel filtering pipeline. We aim to take a step towards alleviating the low resource status of the Bangla languag…

DiversityVocal Bursts Intensity Prediction

BanglaNirTox: A Large-scale Parallel Corpus for Explainable AI in Bengali Text Detoxification

2025-11-03 · Ayesha Afroza Mohsin, Mashrur Ahsan, Nafisa Maliyat, Shanta Maria 외 arxiv

Toxic language in Bengali remains prevalent, especially in online environments, with few effective precautions against it. Although text detoxification has seen progress in high-resource languages, Bengali remains undere…

Preparation of Bangla Speech Corpus from Publicly Available Audio \& Text

2020-05-01 · LREC 2020 5 · Shafayat Ahmed, Nafis Sadeq, Sudipta Saha Shubha, Md. Nahidul Islam 외

Automatic speech recognition systems require large annotated speech corpus. The manual annotation of a large corpus is very difficult. In this paper, we focus on the automatic preparation of a speech corpus for Banglades…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+2

Corpus-Based Paraphrase Detection Experiments and Review

2021-05-31 · Tedo Vrbanec, Ana Mestrovic

Paraphrase detection is important for a number of applications, including plagiarism detection, authorship attribution, question answering, text summarization, text mining in general, etc. In this paper, we give a perfor…

Authorship AttributionDeep LearningModel SelectionQuestion Answering+3

BanglaEmbed: Efficient Sentence Embedding Models for a Low-Resource Language Using Cross-Lingual Distillation Techniques

2024-11-22 · Muhammad Rafsan Kabir, Md. Mohibur Rahman Nabil, Mohammad Ashrafuzzaman Khan

Sentence-level embedding is essential for various tasks that require understanding natural language. Many studies have explored such embeddings for high-resource languages like English. However, low-resource languages li…

Hate Speech DetectionKnowledge DistillationSemantic Textual SimilaritySentence+3