A 200K Malware Hash Dataset with Experimental Splits for Machine Learning Classification

Published: 14 September 2026| Version 1 | DOI: 10.17632/xv84w8cc72.1
Contributors:
Antonio Lara Gutiérrez, Nicolas Alvarez,

Description

This dataset provides a comprehensive manifest of 201,549 SHA-256 hashes representing both malicious (malware) and benign (goodware) software samples. It is specifically designed to facilitate reproducibility in machine learning and artificial intelligence experiments for malware detection. The provided manifest.csv file includes the overarching binary labels as well as granular categorizations (e.g., Trojan, Obfuscated). Crucially, it contains pre-defined train/test splits and boolean masks (in_binary_experiment, in_multiclass_experiment, in_trimodal_binary, in_trimodal_multiclass) that dictate the exact subset of samples utilized across different experimental configurations. This structure allows researchers to perfectly replicate the dataset splits used in the associated study.

Files

Institutions

Categories

Cybersecurity, Meta Dataset

Licence