paper-with-me

홈 › Papers

CTIBench: A Benchmark for Evaluating LLMs in Cyber Threat Intelligence

2024-06-11 · Md Tanvirul Alam, Dipkamal Bhusal, Le Nguyen, Nidhi Rastogi

Cyber threat intelligence (CTI) is crucial in today's cybersecurity landscape, providing essential insights to understand and mitigate the ever-evolving cyber threats. The recent rise of Large Language Models (LLMs) have shown potential in this domain, but concerns about their reliability, accuracy, and hallucinations persist. While existing benchmarks provide general evaluations of LLMs, there are no benchmarks that address the practical and applied aspects of CTI-specific tasks. To bridge this gap, we introduce CTIBench, a benchmark designed to assess LLMs' performance in CTI applications. CTIBench includes multiple datasets focused on evaluating knowledge acquired by LLMs in the cyber-threat landscape. Our evaluation of several state-of-the-art models on these tasks provides insights into their strengths and weaknesses in CTI contexts, contributing to a better understanding of LLM capabilities in CTI.

📄 PDF Abstract BibTeX arXiv:2406.07599

Code (1)

xashru/cti-bench 공식 구현

Similar Papers 제목 키워드 기반

AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence

2025-11-03 · Md Tanvirul Alam, Dipkamal Bhusal, Salman Ahmad, Nidhi Rastogi 외 arxiv

Large Language Models (LLMs) have demonstrated strong capabilities in natural language reasoning, yet their application to Cyber Threat Intelligence (CTI) remains limited. CTI analysis involves distilling large volumes o…

OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities

2025-02-18 · Michael Kouremetis, Marissa Dotter, Alex Byrne, Dan Martin 외

The prospect of artificial intelligence (AI) competing in the adversarial landscape of cyber security has long been considered one of the most impactful, challenging, and potentially dangerous applications of AI. Here, w…

Large Language ModelMultiple-choice

PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities

2025-10-13 · Zicheng Liu, Lige Huang, Jie Zhang, Dongrui Liu 외 arxiv

The increasing autonomy of Large Language Models (LLMs) necessitates a rigorous evaluation of their potential to aid in cyber offense. Existing benchmarks often lack real-world complexity and are thus unable to accuratel…

Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting

2025-09-28 · Yuqiao Meng, Luoxi Tang, Feiyang Yu, Xi Li 외 arxiv

As cyber threats continue to grow in scale and sophistication, blue team defenders increasingly require advanced tools to proactively detect and mitigate risks. Large Language Models (LLMs) offer promising capabilities f…

CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge

2024-02-12 · Norbert Tihanyi, Mohamed Amine Ferrag, Ridhi Jain, Tamas Bisztray 외

Large Language Models (LLMs) are increasingly used across various domains, from software development to cyber threat intelligence. Understanding all the different fields of cybersecurity, which includes topics such as cr…

General KnowledgeMultiple-choiceRAGRetrieval+1