Harmonized Multi-Source Fraud Transaction Dataset for Benchmarking Synthetic Data Generation Methods

Published: 11 July 2026| Version 2 | DOI: 10.17632/gcsx4mfzhb.2
Contributors:
Monica a,
,

Description

This repository contains a harmonized multi-source fraud transaction dataset created by integrating records from the IEEE-CIS Fraud Detection, PaySim, and Credit Card Fraud Detection datasets into a unified schema. The harmonized dataset consists of 5,000 transaction records with balanced class labels (2,500 legitimate and 2,500 fraudulent transactions) and includes common features such as transaction time, amount, transaction type, fraud label, and source dataset. In addition to the harmonized dataset, synthetic benchmark datasets generated using SMOTE, CTGAN, and WGAN are provided to support comparative evaluation of synthetic data generation methods. A data dictionary and feature harmonization specification are also included. This dataset is intended to facilitate reproducible research in fraud detection, synthetic data generation, and machine learning benchmarking and accompanies the related Data in Brief article.

Files

Steps to reproduce

Download the harmonized dataset and supporting files from this repository. Refer to the data dictionary and feature harmonization specification to understand the variables. Use the harmonized dataset (harmonized_fraud_dataset_5000.csv) for fraud detection experiments. Use the provided SMOTE, CTGAN, and WGAN datasets to benchmark synthetic data generation methods. Compare model performance using standard evaluation metrics such as Accuracy, Precision, Recall, F1-score, ROC-AUC, and PR-AUC as described in the accompanying Data in Brief article.

Institutions

Categories

Computer Science, Artificial Intelligence, Machine Learning

Licence