An Annotated Image Dataset for Cucumber Leaf Disease Detection and Segmentation in Complex Orchard Environments
Description
This dataset was developed to support automated detection and localization of cucumber leaf abnormalities under complex orchard conditions. The research hypothesis is that visually observable disease damage and abnormal leaf coloration can be identified from field images using deep-learning-based classification and object-detection models, despite variations in illumination, viewing angle, scale, background clutter, leaf overlap, and growth stage. The dataset contains 500 cucumber leaf images: 250 images of healthy leaves and 250 images showing disease symptoms or abnormal coloration. Images were acquired in real cultivation environments and subsequently inspected and organized according to leaf condition. The annotated subset is accompanied by text label files containing rectangular bounding-box annotations. Each annotation row has five values: the class identifier, the horizontal and vertical coordinates of the bounding-box center, and the box width and height. Class 0 represents abnormal color—that is, a visible leaf color other than healthy green, such as yellowing or browning—whereas class 1 represents visible disease-related damage. An image may contain multiple annotations because several affected regions can occur on the same leaf. The dataset includes 2,705 annotated regions: 1,741 color regions and 964 damage regions. The annotation distribution indicates that abnormal coloration occurs more frequently than directly identifiable damage regions. This reflects the heterogeneous visual presentation of cucumber leaf disorders and makes the dataset useful for studying both general discoloration and localized physical symptoms. The healthy images can support image-level binary classification, while the annotated images can be used for object detection and rectangular region localization. The current dataset version provides bounding boxes rather than pixel-level segmentation masks; therefore, it should not be interpreted as conventional semantic- or instance-segmentation ground truth. Researchers may use the dataset to train, validate, and compare deep-learning architectures for healthy-versus-affected leaf classification, color-abnormality detection, disease-damage localization, transfer learning, data augmentation, and robustness evaluation under natural cultivation conditions. Before model development, users should verify image–label correspondence, preserve the class mapping, and create non-overlapping training, validation, and test partitions. Because the annotations describe visible phenotypic categories rather than pathogen-specific diagnoses, predictions should be interpreted as indications of abnormal coloration or damage and not as confirmation of a particular cucumber disease.
Files
Steps to reproduce
1. Select cucumber-growing environments containing healthy plants and plants with visible leaf abnormalities. Include different growth stages, symptom severities, viewpoints, lighting conditions, and backgrounds. 2. Photograph the cucumber leaves using a smartphone or comparable digital camera. Capture images from different angles and distances under natural cultivation conditions, including complex backgrounds, overlapping leaves, shadows, and variable illumination. 3. Visually inspect and organize the images into two groups: healthy leaves and leaves with disease-related damage or abnormal coloration. The final dataset contains 500 images: 250 healthy-leaf images and 250 affected-leaf images. 4. Perform quality control by removing corrupted, duplicate, severely blurred, overexposed, or underexposed images. Retain images in which the leaves and relevant symptoms can be interpreted. 5. Define two annotation classes. Class 0 (color) represents visible deviations from healthy green coloration, such as yellow or brown regions. Class 1 (damage) represents visible disease-related or physically damaged leaf regions. The labels describe observable symptoms and not pathogen-confirmed diagnoses. 6. Import the affected-leaf images into an annotation platform such as Label Studio. Draw rectangular bounding boxes around all visible color-abnormality and damage regions. Assign the appropriate class to each box. One image may contain multiple annotations. 7. Export one TXT label file for each annotated image. Every annotation row contains five values in the following order: class identifier, bounding-box center coordinate x, center coordinate y, box width, and box height. Preserve and document the coordinate-normalization convention used during export. 8. Verify the annotations by checking image–label correspondence, class assignments, box placement, row structure, and allowed class identifiers. The completed dataset contains 2,705 bounding boxes: 1,741 color annotations and 964 damage annotations. 9. Organize the images and labels in clearly named folders and provide a README describing the classes and annotation format. Before model development, divide the data into non-overlapping training, validation, and test sets to prevent data leakage. 10. The healthy and affected images may be used for image classification, while the bounding boxes support object detection and region localization. The current version does not contain pixel-level segmentation masks; therefore, any masks derived from bounding boxes should be described as approximate rather than expert-annotated segmentation ground truth.
Institutions
- Joldasbekov Institute of Mechanics and EngineeringAlmaty, Almaty