Kurdistan Watermelon Disease Classification Image Dataset (Halabja, Penjwin, and Sharazur)
Description
The Kurdistan Watermelon Disease Classification Image Dataset (Halabja, Penjwin, and Sharazur) is a curated collection of watermelon images acquired from agricultural fields in the Halabja, Penjwin, and Sharazur districts of the Kurdistan Region, Iraq, an important watermelon-producing area. The dataset contains 1,152 original high-resolution images categorized into five classes: Healthy, Anthracnose, Downy Mildew, Iron Deficiency Chlorosis, and Spider Mite Infestation. To enhance dataset diversity and improve the robustness and generalization of artificial intelligence models, 10,368 augmented images were generated using a variety of image augmentation techniques, including horizontal flipping, random rotation, Gaussian blur, Gaussian noise, cropping, perspective transformation, contrast adjustment, and brightness variation. All images were resized to 512 × 512 pixels to ensure consistency and compatibility with modern deep learning frameworks. The images were captured under uncontrolled field conditions with natural illumination, reflecting realistic agricultural environments and varying backgrounds. This dataset is intended to support research on automated watermelon disease classification using computer vision and artificial intelligence techniques, while also serving as a valuable resource for developing robust deep learning models for smart agriculture applications. 1_Original_Captured_Images contains the raw images exactly as captured in the field using mobile devices. These images retain their original resolutions, formats (e.g., JPG, jpg, HEIC), and filenames without any preprocessing. 2_Standardized_Dataset contains the preprocessed version of the dataset, where all images have been converted to PNG format, resized to 512 × 512 pixels, and renamed using a consistent class-based naming convention. This folder is intended to serve as the primary benchmark dataset for evaluation and comparison of machine learning and deep learning models. 3_Augmented_Dataset contains the complete training dataset, including all standardized original images together with their augmented versions. The augmented images were generated using realistic image augmentation techniques such as horizontal flipping, random rotation, Gaussian blur, Gaussian noise, cropping, perspective transformation, contrast adjustment, and brightness variation to increase data diversity and improve model robustness. The original standardized images are also included in this folder, allowing researchers to use it directly as a ready-to-train dataset without requiring additional preprocessing or augmentation.