Vulnerability Datasets and Code to improve Severity Classification using length-based imbalance handling
Description
This dataset contains software vulnerabilities collected from the National Vulnerability Database (NVD) for description-based vulnerability severity classification. It includes a heterogeneous Mixed dataset and product-specific Windows and Android datasets. Vulnerabilities reported during 2020–2024 are provided for model development and training, while vulnerabilities reported during 2025–2026 are provided as temporally separated test data. Each record includes the vulnerability description and its CVSS v3 severity class (Low, Medium, High, or Critical). The Windows and Android datasets were extracted using CPE-based filtering. The datasets support experiments on class imbalance handling, temporal generalization, and generalist versus specialist vulnerability classification. The code files contain the implementation used for the experiments reported in the study. They include preprocessing and feature representation, the proposed length-based imbalance handling method, conventional imbalance handling techniques, and the training and evaluation of SVM, Random Forest, CNN, LSTM, and DistilBERT models. The code also covers temporal evaluation on future vulnerabilities and the comparison of generalist and specialist models on Windows and Android datasets. Evaluation includes accuracy, balanced accuracy, macro-F1, class-wise performance, and MCC.
Files
Institutions
- Netaji Subhas University of TechnologyDelhi, New Delhi