BnPC: A Corpus for Paraphrase Detection in Bangla
In this paper, we present the first benchmark dataset for paraphrase detection in Bangla language. Despite being the sixth most spoken language in the world, paraphrase identification in the Bangla language is barely explored. Our dataset contains 8,787 human-annotated sentence pairs collected from a total of 23 newspaper outlets' headlines on four categories. We explore different linguistic features and pre-trained language models to benchmark the dataset. We perform a human evaluation experiment to obtain a better understanding of the task's constraints, which reveals intriguing insights. We make our dataset and code publicly available.
Code (0)
등록된 구현이 없습니다.
Tasks
Paraphrase IdentificationSentenceSimilar Papers 제목 키워드 기반
BanglaParaphrase: A High-Quality Bangla Paraphrase Dataset
In this work, we present BanglaParaphrase, a high-quality synthetic Bangla Paraphrase dataset curated by a novel filtering pipeline. We aim to take a step towards alleviating the low resource status of the Bangla languag…
DiversityVocal Bursts Intensity PredictionBanglaNirTox: A Large-scale Parallel Corpus for Explainable AI in Bengali Text Detoxification
Toxic language in Bengali remains prevalent, especially in online environments, with few effective precautions against it. Although text detoxification has seen progress in high-resource languages, Bengali remains undere…
Preparation of Bangla Speech Corpus from Publicly Available Audio \& Text
Automatic speech recognition systems require large annotated speech corpus. The manual annotation of a large corpus is very difficult. In this paper, we focus on the automatic preparation of a speech corpus for Banglades…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+2Corpus-Based Paraphrase Detection Experiments and Review
Paraphrase detection is important for a number of applications, including plagiarism detection, authorship attribution, question answering, text summarization, text mining in general, etc. In this paper, we give a perfor…
Authorship AttributionDeep LearningModel SelectionQuestion Answering+3BanglaEmbed: Efficient Sentence Embedding Models for a Low-Resource Language Using Cross-Lingual Distillation Techniques
Sentence-level embedding is essential for various tasks that require understanding natural language. Many studies have explored such embeddings for high-resource languages like English. However, low-resource languages li…
Hate Speech DetectionKnowledge DistillationSemantic Textual SimilaritySentence+3