paper-with-me

Papers

Primus: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM Training

2025-02-16 · Yao-Ching Yu, Tsun-Han Chiang, Cheng-Wei Tsai, Chien-Ming Huang, Wen-Kwang Tsao

Large Language Models (LLMs) have shown remarkable advancements in specialized fields such as finance, law, and medicine. However, in cybersecurity, we have noticed a lack of open-source datasets, with a particular lack of high-quality cybersecurity pretraining corpora, even though much research indicates that LLMs acquire their knowledge during pretraining. To address this, we present a comprehensive suite of datasets covering all major training stages, including pretraining, instruction fine-tuning, and reasoning distillation with cybersecurity-specific self-reflection data. Extensive ablation studies demonstrate their effectiveness on public cybersecurity benchmarks. In particular, continual pre-training on our dataset yields a 15.88% improvement in the aggregate score, while reasoning distillation leads to a 10% gain in security certification (CISSP). We will release all datasets and trained cybersecurity LLMs under the ODC-BY and MIT licenses to encourage further research in the community. For access to all datasets and model weights, please refer to https://huggingface.co/collections/trendmicro-ailab/primus-67b1fd27052b802b4af9d243.

📄 PDF Abstract BibTeX arXiv:2502.11191

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PRIMUS: Pretraining IMU Encoders with Multimodal Self-Supervision

2024-11-22 · Arnav M. Das, Chi Ian Tang, Fahim Kawsar, Mohammad Malekzadeh

Sensing human motions through Inertial Measurement Units (IMUs) embedded in personal devices has enabled significant applications in health and wellness. Labeled IMU data is scarce, however, unlabeled or weakly labeled I…

Primus: Enforcing Attention Usage for 3D Medical Image Segmentation

2025-03-03 · Tassilo Wald, Saikat Roy, Fabian Isensee, Constantin Ulrich 외

Transformers have achieved remarkable success across multiple fields, yet their impact on 3D medical image segmentation remains limited with convolutional networks still dominating major benchmarks. In this work, we a) a…

Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation

A High-Accuracy Optical Music Recognition Method Based on Bottleneck Residual Convolutions

2026-04-07 · Junwen Ma, Huhu Xue, Xingyuan Zhao, and Weicheng Fu arxiv

Optical Music Recognition (OMR) aims to convert printed or handwritten music score images into editable symbolic representations. This paper presents an end-to-end OMR framework that combines residual bottleneck convolut…

Computational Efficiency

ActPC-Chem: Discrete Active Predictive Coding for Goal-Guided Algorithmic Chemistry as a Potential Cognitive Kernel for Hyperon & PRIMUS-Based AGI

2024-12-21 · Ben Goertzel

We explore a novel paradigm (labeled ActPC-Chem) for biologically inspired, goal-guided artificial intelligence (AI) centered on a form of Discrete Active Predictive Coding (ActPC) operating within an algorithmic chemist…

Aqulia-Med LLM: Pioneering Full-Process Open-Source Medical Language Models

2024-06-18 · Lulu Zhao, Weihao Zeng, Xiaofeng Shi, Hua Zhou 외

Recently, both closed-source LLMs and open-source communities have made significant strides, outperforming humans in various general domains. However, their performance in specific professional fields such as medicine, e…

Multiple-choice