paper-with-me

Papers

NBC-Softmax : Darkweb Author fingerprinting and migration tracking

2022-12-15 · Gayan K. Kulatilleke, Shekhar S. Chandra, Marius Portmann

Metric learning aims to learn distances from the data, which enhances the performance of similarity-based algorithms. An author style detection task is a metric learning problem, where learning style features with small intra-class variations and larger inter-class differences is of great importance to achieve better performance. Recently, metric learning based on softmax loss has been used successfully for style detection. While softmax loss can produce separable representations, its discriminative power is relatively poor. In this work, we propose NBC-Softmax, a contrastive loss based clustering technique for softmax loss, which is more intuitive and able to achieve superior performance. Our technique meets the criterion for larger number of samples, thus achieving block contrastiveness, which is proven to outperform pair-wise losses. It uses mini-batch sampling effectively and is scalable. Experiments on 4 darkweb social forums, with NBCSAuthor that uses the proposed NBC-Softmax for author and sybil detection, shows that our negative block contrastive approach constantly outperforms state-of-the-art methods using the same network architecture. Our code is publicly available at : https://github.com/gayanku/NBC-Softmax

📄 PDF Abstract BibTeX arXiv:2212.08184

Code (4)

gayanku/nbc-softmax 공식 구현 pytorch
MindCode-4/code-8/tree/main/NBC-Softmax mindspore
MindSpore-scientific/code-6/tree/main/NBC-Softmax mindspore
nanzhaogang/contrib/tree/master/application/NBC-Softmax mindspore

Tasks

Metric LearningStyle Detection

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

VeriDark: A Large-Scale Benchmark for Authorship Verification on the Dark Web

2022-07-07 · Andrei Manolache, Florin Brad, Antonio Barbalau, Radu Tudor Ionescu 외

The DarkWeb represents a hotbed for illicit activity, where users communicate on different market forums in order to exchange goods and services. Law enforcement agencies benefit from forensic tools that perform authorsh…

Authorship Verification

A Behavioral Fingerprint for Large Language Models: Provenance Tracking via Refusal Vectors

2026-02-10 · Zhenyu Xu, Victor S. Sheng arxiv

Protecting the intellectual property of large language models (LLMs) is a critical challenge due to the proliferation of unauthorized derivative models. We introduce a novel fingerprinting framework that leverages the be…

Fingerprinting and Tracing Shadows: The Development and Impact of Browser Fingerprinting on Digital Privacy

2024-11-18 · Alexander Lawall

Browser fingerprinting is a growing technique for identifying and tracking users online without traditional methods like cookies. This paper gives an overview by examining the various fingerprinting techniques and analyz…

Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets

2026-05-27 · Andrea Gurioli, Davide D'Ascenzo, Federico Pennino, Maurizio Gabbrielli 외 arxiv

Large language models (LLMs) for code completion and generation are increasingly used in software development, yet they may reproduce training examples verbatim and without authorship attribution, raising legal and ethic…

Code Completion

Determinants Of Migration: Linear regression Analysis in Indian Context

2022-09-23 · Soumik Ghosh, Arpan Chakraborty

Environmental degradation, global pandemic and severing natural resource related problems cater to increase demand resulting from migration is nightmare for all of us. Huge flocks of people are rushing towards to earn, t…

regression