Formation of a General Training Dataset for Breast Cancer Based on Nominal Features

Published: 15 April 2026| Version 1 | DOI: 10.17632/fsb6wdyzpy.1
Contributors:
,
,
,
,

Description

This dataset was developed based on clinical data collected from the Surkhandarya branch of the Republican Specialized Scientific and Practical Medical Center of Oncology and Radiology of the Republic of Uzbekistan. During the study, breast cancer patient records were thoroughly analyzed in collaboration with oncologists and domain experts, and key clinical and anamnestic features (symptoms) were identified.Initially, a total of 1,054 patient records were collected for the study. During the preprocessing stage, the records were evaluated for completeness, reliability, and consistency with the selected features. Records that lacked sufficient information in sections such as complaints, medical history, life history, epidemiological history, local status, and major physiological systems (respiratory, cardiovascular, digestive, and urinary systems) were considered unsuitable and excluded from the dataset. As a result, a final training dataset consisting of 567 complete and reliable instances was formed. The dataset is represented in a nominal feature space, where each instance is described using 32 clinical and diagnostic features. These features were identified in collaboration with medical experts and represent key indicators for the early detection of breast cancer. At the next stage, an informative feature selection algorithm was applied to identify the most significant features, resulting in a reduced set of 18 features. Based on these selected features, the dataset was classified into 13 distinct classes. The distribution of instances across classes is as follows: Class 1 – 54 instances, Class 2 – 73, Class 3 – 19, Class 4 – 39, Class 5 – 16, Class 6 – 5, Class 7 – 276, Class 8 – 45, Class 9 – 8, Class 10 – 10, Class 11 – 10, Class 12 – 7, and Class 13 – 5 instances. This dataset is intended for use in early diagnosis, classification, and predictive modeling of breast cancer, and can support the development of machine learning algorithms as well as clinical decision-making systems.

Files

Steps to reproduce

This dataset was collected from clinical records of breast cancer patients at the Surkhandarya branch of the Republican Specialized Scientific and Practical Medical Center of Oncology and Radiology. The data collection process was carried out in collaboration with oncologists and medical experts. A total of 1,054 patient records were initially gathered. During preprocessing, the data were carefully reviewed for completeness, consistency, and relevance to the selected features. Records with missing or incomplete information in key sections such as patient complaints, medical history, life history, epidemiological history, local examination, and major physiological systems (respiratory, cardiovascular, digestive, and urinary systems) were excluded. As a result, 567 complete and reliable records were selected to form the final dataset. Each instance in the dataset is represented using 32 nominal clinical features describing patient symptoms and conditions. The dataset was encoded using categorical (nominal) values to facilitate computational processing. Additionally, feature selection techniques were applied to identify the most informative features, improving the quality of the dataset for machine learning applications.

Categories

Oncology, Breast Cancer, Machine Learning

Licence