paper-with-me

Papers

Using Machine Learning to Detect Fraudulent SMSs in Chichewa

2025-02-24 · Amelia Taylor, Amoss Robert

SMS enabled fraud is of great concern globally. Building classifiers based on machine learning for SMS fraud requires the use of suitable datasets for model training and validation. Most research has centred on the use of datasets of SMSs in English. This paper introduces a first dataset for SMS fraud detection in Chichewa, a major language in Africa, and reports on experiments with machine learning algorithms for classifying SMSs in Chichewa as fraud or non-fraud. We answer the broader research question of how feasible it is to develop machine learning classification models for Chichewa SMSs. To do that, we created three datasets. A small dataset of SMS in Chichewa was collected through primary research from a segment of the young population. We applied a label-preserving text transformations to increase its size. The enlarged dataset was translated into English using two approaches: human translation and machine translation. The Chichewa and the translated datasets were subjected to machine classification using random forest and logistic regression. Our findings indicate that both models achieved a promising accuracy of over 96% on the Chichewa dataset. There was a drop in performance when moving from the Chichewa to the translated dataset. This highlights the importance of data preprocessing, especially in multilingual or cross-lingual NLP tasks, and shows the challenges of relying on machine-translated text for training machine learning models. Our results underscore the importance of developing language specific models for SMS fraud detection to optimise accuracy and performance. Since most machine learning models require data preprocessing, it is essential to investigate the impact of the reliance on English-specific tools for data preprocessing.

📄 PDF Abstract BibTeX arXiv:2502.16947

Code (0)

등록된 구현이 없습니다.

Tasks

Fraud DetectionMachine Translation

Similar Papers 제목 키워드 기반

Spam Detection Using BERT

2022-06-06 · Thaer Sahmoud, Dr. Mohammad Mikki

Emails and SMSs are the most popular tools in today communications, and as the increase of emails and SMSs users are increase, the number of spams is also increases. Spam is any kind of unwanted, unsolicited digital comm…

Spam detection

SMSSVD - SubMatrix Selection Singular Value Decomposition

2017-10-23 · Rasmus Henningsson, Magnus Fontes

High throughput biomedical measurements normally capture multiple overlaid biologically relevant signals and often also signals representing different types of technical artefacts like e.g. batch effects. Signal identifi…

Dimensionality Reduction

Detection of fraudulent users in P2P financial market

2019-09-24 · Hao Wang

Financial fraud detection is one of the core technological assets of Fintech companies. It saves tens of millions of money fro m Chinese Fintech companies since the bad loan rate is more than 10%. HC Financial Service Gr…

BIG-bench Machine LearningFraud Detection

PrismSSL: One Interface, Many Modalities; A Single-Interface Library for Multimodal Self-Supervised Learning

2025-11-21 · Melika Shirian, Kianoosh Vadaei, Kian Majlessi, Audrina Ebrahimi 외 arxiv

We present PrismSSL, a Python library that unifies state-of-the-art self-supervised learning (SSL) methods across audio, vision, graphs, and cross-modal settings in a single, modular codebase. The goal of the demo is to …

Self-Supervised Learning

Performance Bounds for Near-Field Localization with Widely-Spaced Multi-Subarray mmWave/THz MIMO

2023-09-12 · Songjie Yang, Xinyi Chen, Yue Xiu, Wanting Lyu 외

This paper investigates the potential of near-field localization using widely-spaced multi-subarrays (WSMSs) and analyzing the corresponding angle and range Cram\'er-Rao bounds (CRBs). By employing the Riemann sum, close…

Form