Sugarcane Disease Image Dataset
Description
It was developed to support research related to the identification of diseases in the leaves of the sugarcane plant using image-based machine learning and deep learning techniques. The hypothesis is that specific features causing distinguishable variations in the color and texture patterns of leaves, along with structural deformity, are sufficient for models to accurately classify diseases. There are three classes in this dataset: Healthy, Red Rot, and White Leaf. Healthy leaves appear uniformly green without visible damage. Red Rot images show reddish-brown lesions and discoloration that are usually the result of fungal infection. White Leaf samples display pale or whitish chlorotic leaves associated with phytoplasma infection. These clear visual differences illustrate the point that disease types can be diagnosed through image analysis and form a basis for the hypothesis that such patterns may effectively be learned by machine learning models. The images were collected directly from the sugarcane fields of Maharashtra, India, using a mobile phone camera under natural daylight. Each image is manually reviewed to ensure its clarity before labeling according to references in agricultural diseases. Natural variation of field conditions-such as lighting conditions, leaf angles, backgrounds, and disease severity-has been captured in the dataset. The dataset can be interpreted by researchers as a practical, field-collected resource suitable for training classification models, analyzing visual disease characteristics, and developing automated crop monitoring systems. This dataset can be used for CNN training, transfer learning, studies on image segmentation, or any plant disease detection pipeline. Users may apply some preprocessing steps such as resizing or normalization before training a model. Overall, the dataset forms a very reliable base for computer vision–based agricultural disease research.
Files
Steps to reproduce
1. Visit sugarcane fields during normal daylight conditions and select leaves representing Healthy, Red Rot, and White Leaf categories. 2. Capture images using a mobile phone camera from a close distance, ensuring clear visibility of leaf texture and disease symptoms. 3. Collect multiple samples for each class from different plants to include natural variation in lighting and appearance. 4. Transfer all captured images to a computer and manually inspect them for clarity, discarding blurred or unusable photos. 5. Create three folders named healthy/, red_rot/, and white_leaf/. 6. Manually label each image by placing it into the corresponding folder based on visible disease characteristics, using agricultural reference guides for confirmation. 7. Compress the folder structure into a single ZIP file to create the final dataset.
Institutions
- Vishwakarma Institute of Information Technology