Authorship Attribution in Bangla literature using Character-level CNN
Characters are the smallest unit of text that can extract stylometric signals to determine the author of a text. In this paper, we investigate the effectiveness of character-level signals in Authorship Attribution of Bangla Literature and show that the results are promising but improvable. The time and memory efficiency of the proposed model is much higher than the word level counterparts but accuracy is 2-5% less than the best performing word-level models. Comparison of various word-based models is performed and shown that the proposed model performs increasingly better with larger datasets. We also analyze the effect of pre-training character embedding of diverse Bangla character set in authorship attribution. It is seen that the performance is improved by up to 10% on pre-training. We used 2 datasets from 6 to 14 authors, balancing them before training and compare the results.
Code (1)
Tasks
Authorship AttributionSimilar Papers 제목 키워드 기반
Authorship Attribution in Bangla Literature (AABL) via Transfer Learning using ULMFiT
Authorship Attribution is the task of creating an appropriate characterization of text that captures the authors' writing style to identify the original author of a given piece of text. With increased anonymity on the in…
Authorship AttributionSentenceTransfer LearningBARD10: A New Benchmark Reveals Significance of Bangla Stop-Words in Authorship Attribution
This research presents a comprehensive investigation into Bangla authorship attribution, introducing a new balanced benchmark corpus BARD10 (Bangla Authorship Recognition Dataset of 10 authors) and systematically analyzi…
Character-level and Multi-channel Convolutional Neural Networks for Large-scale Authorship Attribution
Convolutional neural networks (CNNs) have demonstrated superior capability for extracting information from raw signals in computer vision. Recently, character-level and multi-channel CNNs have exhibited excellent perform…
Authorship AttributionSentenceSentence ClassificationIntegrating Bidirectional Long Short-Term Memory with Subword Embedding for Authorship Attribution
The problem of unveiling the author of a given text document from multiple candidate authors is called authorship attribution. Manifold word-based stylistic markers have been successfully used in deep learning methods to…
Authorship AttributionCCTAA: A Reproducible Corpus for Chinese Authorship Attribution Research
Authorship attribution infers the likely author of an unsigned, single-authored document from a pool of candidates. Despite recent advances, a lack of standard, reproducible testbeds for Chinese language documents impede…
Authorship Attribution