paper-with-me

홈 › Papers

Automated Classification of Cybercrime Complaints using Transformer-based Language Models for Hinglish Texts

2024-12-21 · Nanda Rani, Divyanshu Singh, Bikash Saha, Sandeep Kumar Shukla

The rise in cybercrime and the complexity of multilingual and code-mixed complaints present significant challenges for law enforcement and cybersecurity agencies. These organizations need automated, scalable methods to identify crime types, enabling efficient processing and prioritization of large complaint volumes. Manual triaging is inefficient, and traditional machine learning methods fail to capture the semantic and contextual nuances of textual cybercrime complaints. Moreover, the lack of publicly available datasets and privacy concerns hinder the research to present robust solutions. To address these challenges, we propose a framework for automated cybercrime complaint classification. The framework leverages Hinglish-adapted transformers, such as HingBERT and HingRoBERTa, to handle code-mixed inputs effectively. We employ the real-world dataset provided by Indian Cybercrime Coordination Centre (I4C) during CyberGuard AI Hackathon 2024. We employ GenAI open source model-based data augmentation method to address class imbalance. We also employ privacy-aware preprocessing to ensure compliance with ethical standards while maintaining data integrity. Our solution achieves significant performance improvements, with HingRoBERTa attaining an accuracy of 74.41% and an F1-score of 71.49%. We also develop ready-to-use tool by integrating Django REST backend with a modern frontend. The developed tool is scalable and ready for real-world deployment in platforms like the National Cyber Crime Reporting Portal. This work bridges critical gaps in cybercrime complaint management, offering a scalable, privacy-conscious, and adaptable solution for modern cybersecurity challenges.

📄 PDF Abstract BibTeX arXiv:2412.16614

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

BEACON: A Unified Behavioral-Tactical Framework for Explainable Cybercrime Analysis with Large Language Models

2025-12-06 · Arush Sachdeva, Rajendraprasad Saravanan, Gargi Sarkar, Kavita Vemuri 외 arxiv

Cybercrime increasingly exploits human cognitive biases in addition to technical vulnerabilities, yet most existing analytical frameworks focus primarily on operational aspects and overlook psychological manipulation. Th…

Multi-Label Classification

Two-step Automated Cybercrime Coded Word Detection using Multi-level Representation Learning

2024-03-16 · Yongyeon Kim, Byung-Won On, Ingyu Lee

In social network service platforms, crime suspects are likely to use cybercrime coded words for communication by adding criminal meanings to existing words or replacing them with similar words. For instance, the word 'i…

Representation Learning

AI Appeals Processor: A Deep Learning Approach to Automated Classification of Citizen Appeals in Government Services

2026-04-04 · Vladimir Beskorovainyi arxiv

Government agencies worldwide face growing volumes of citizen appeals, with electronic submissions increasing significantly over recent years. Traditional manual processing averages 20 minutes per appeal with only 67% cl…

Computational Efficiency

Evaluating robustness of language models for chief complaint extraction from patient-generated text

2019-11-15 · Ilya Valmianski, Caleb Goodwin, Ian M. Finn, Naqi Khan 외

Automated classification of chief complaints from patient-generated text is a critical first step in developing scalable platforms to triage patients without human intervention. In this work, we evaluate several approach…

Modeling the Severity of Complaints in Social Media

2021-03-23 · NAACL 2021 4 · Mali Jin, Nikolaos Aletras

The speech act of complaining is used by humans to communicate a negative mismatch between reality and expectations as a reaction to an unfavorable situation. Linguistic theory of pragmatics categorizes complaints into v…