paper-with-me

Papers

Classifying Illegal Activities on Tor Network Based on Web Textual Contents

2017-04-01 · EACL 2017 4 · Mhd Wesam Al Nabki, Eduardo Fidalgo, Enrique Alegre, Ivan de Paz

The freedom of the Deep Web offers a safe place where people can express themselves anonymously but they also can conduct illegal activities. In this paper, we present and make publicly available a new dataset for Darknet active domains, which we call {''}Darknet Usage Text Addresses{''} (DUTA). We built DUTA by sampling the Tor network during two months and manually labeled each address into 26 classes. Using DUTA, we conducted a comparison between two well-known text representation techniques crossed by three different supervised classifiers to categorize the Tor hidden services. We also fixed the pipeline elements and identified the aspects that have a critical influence on the classification results. We found that the combination of TFIDF words representation with Logistic Regression classifier achieves 96.6{\%} of 10 folds cross-validation accuracy and a macro F1 score of 93.7{\%} when classifying a subset of illegal activities from DUTA. The good performance of the classifier might support potential tools to help the authorities in the detection of these activities.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Logistic Regression Logistic Regression, despite its name, is a linear model for classification rather than regression. Logistic regression is also known in the literature as logit regression,…

Similar Papers 제목 키워드 기반

Illicit Darkweb Classification via Natural-language Processing: Classifying Illicit Content of Webpages based on Textual Information

2023-12-08 · Giuseppe Cascavilla, Gemma Catolino, Mirella Sangiovanni

This work aims at expanding previous works done in the context of illegal activities classification, performing three different steps. First, we created a heterogeneous dataset of 113995 onion sites and dark marketplaces…

ClassificationLanguage ModelingLanguage Modellingtext-classification+1

When the Few Outweigh the Many: Illicit Content Recognition with Few-Shot Learning

2023-11-28 · G. Cascavilla, G. Catolino, M. Conti, D. Mellios 외

The anonymity and untraceability benefits of the Dark web account for the exponentially-increased potential of its popularity while creating a suitable womb for many illicit activities, to date. Hence, in collaboration w…

Few-Shot Learning

FreezeAsGuard: Mitigating Illegal Adaptation of Diffusion Models via Selective Tensor Freezing

2024-05-24 · Kai Huang, Haoming Wang, Wei Gao

Text-to-image diffusion models can be fine-tuned in custom domains to adapt to specific user preferences, but such adaptability has also been utilized for illegal purposes, such as forging public figures' portraits, dupl…

A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities

2026-05-28 · Kenji Imamura, Masao Ideuchi, Atsushi Fujita arxiv

In this paper, we discuss question-answer dataset for LLM safety evaluation, with a focus on illegal activities. Specifically, on the basis of manual analysis of AnswerCarefully, we introduce several additional informati…

Composite Event Recognition for Maritime Monitoring

2019-03-07 · Manolis Pitsikalis, Alexander Artikis, Richard Dreo, Cyril Ray 외

Maritime monitoring systems support safe shipping as they allow for the real-time detection of dangerous, suspicious and illegal vessel activities. We present such a system using the Run-Time Event Calculus, a composite …

Computational EfficiencyPosition