paper-with-me

홈 › Papers

ARVO: Atlas of Reproducible Vulnerabilities for Open-Source Software

2026-06-15 · Xiang Mei, Jordi Del Castillo, Pulkit Singh Singaria, Haoran Xi, Abdelouahab Benchikh, Tiffany Bao, Ruoyu Wang, Yan Shoshitaishvili, Adam Doupé, Hammond Pearce, Brendan Dolan-Gavitt arxiv

Achieving reproducibility, quantity, and diversity in vulnerability datasets has long been viewed as an inherent three-way trade-off, where improving one dimension often comes at the cost of the others. In practice, reproducibility has been the dimension most often neglected. This has limited what can be automatically extracted from historical bug datasets, and has reduced their utility for downstream security research. In this work, we propose a method to produce a new security dataset which ensures reproducibility for diverse vulnerabilities at scale by identifying the key obstacles to large-scale bug reproduction and addressing them with general solutions. Using this method, we introduce full reproducibility to the largest open source software vulnerability dataset (OSS-Fuzz) and construct the ARVO dataset (an Atlas of Reproducible Vulnerabilities in Open-source software). ARVO is a large-scale dataset consisting of over 6,100 real-world vulnerabilities across 311 projects. Focusing on reproducibility, ARVO differs from existing datasets by providing each vulnerability in a form that can be consistently rebuilt, triggered, and analyzed across versions. Reproducibility also enables automatic identification of the corresponding patch for each vulnerability and supports direct interaction with vulnerabilities after code changes, capabilities that existing large-scale datasets do not provide. In our evaluation, ARVO successfully reproduces 81% of vulnerabilities and achieves 89.4% accuracy on the located patches. We also discuss ARVO's influence on both upstream practices and downstream security research.

📄 PDF Abstract BibTeX arXiv:2606.17283

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ARVO: Atlas of Reproducible Vulnerabilities for Open Source Software

2024-08-04 · Xiang Mei, Pulkit Singh Singaria, Jordi Del Castillo, Haoran Xi 외

High-quality datasets of real-world vulnerabilities are enormously valuable for downstream research in software security, but existing datasets are typically small, require extensive manual effort to update, and are miss…

ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis

2025-05-26 · Hawau Olamide Toyin, Rufael Marew, Humaid Alblooshi, Samar M. Magdy 외

We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for other tasks such as speech-based diacri…

DeepFake DetectionFace SwappingSpeech SynthesisVoice Conversion

WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables

2025-12-25 · Zhaojiang Lin, Yong Xu, Kai Sun, Jing Zheng 외 arxiv

Wearable devices such as AI glasses are transforming voice assistants into always-available, hands-free collaborators that integrate seamlessly with daily life, but they also introduce challenges like egocentric audio af…

Marvolo: Programmatic Data Augmentation for Practical ML-Driven Malware Detection

2022-06-07 · Michael D. Wong, Edward Raff, James Holt, Ravi Netravali

Data augmentation has been rare in the cyber security domain due to technical difficulties in altering data in a manner that is semantically consistent with the original data. This shortfall is particularly onerous given…

Data AugmentationMalware Detection

Beyond coauthorship: semantic structure and phantom collaborators in transportation research, 1967--2025

2026-04-26 · Seongjin Choi arxiv

We present a semantic-structural atlas of transportation research built from 120{,}323 papers across 34 peer-reviewed journals published between 1967 and 2025, roughly an order of magnitude larger than and a decade beyon…