Dataset of Common Papaya Diseases: Anthracnose, Mealybug & Whitefly, Mosaic, and Leaf Curl Virus
Description
The Papaya Disease Dataset consists of images and related information on major diseases affecting papaya plants. The dataset is designed to support the development of machine learning and computer vision models for disease identification, classification, and early detection in papaya crops. It includes four main disease categories: Anthracnose of Papaya – A fungal disease causing dark, sunken spots on fruits and leaves. Mealybug and Whitefly Infestation – Caused by insect pests that damage leaves and fruits by sucking sap, leading to curling and yellowing. Mosaic of Papaya Leaf – A viral disease characterized by mottled, patchy patterns of light and dark green areas on leaves. Papaya Leaf Curl Virus – A serious viral disease that leads to upward curling, distortion, and thickening of leaves.
Files
Steps to reproduce
Data Collection Collect papaya leaf and fruit images affected by the four target diseases: Anthracnose of Papaya Mealybug and Whitefly Mosaic of Papaya Leaf Papaya Leaf Curl Virus Capture healthy samples for comparison if needed. Use a good-quality camera or smartphone in natural light to ensure clear images. Data Organization Create four main folders named: Anthracnose Mealybug_Whitefly Mosaic Leaf_Curl_Virus Save each image in its corresponding folder. (Optional) Create a Healthy folder for control samples. Data Preprocessing Resize all images to a uniform dimension (e.g., 224×224 or 256×256 pixels). Normalize pixel values (scale between 0 and 1). Apply data augmentation (rotation, flip, zoom, shift) to increase dataset diversity. Split the dataset into: Training set (70%) Validation set (15%) Testing set (15%) Model Training Load the dataset into a deep learning framework such as TensorFlow, Keras, or PyTorch. Use a Convolutional Neural Network (CNN) model like VGG16, ResNet50, or a custom CNN architecture. Train the model to classify images into the four disease categories. Evaluation Test the trained model using the test dataset. Evaluate performance using metrics such as accuracy, precision, recall, and F1-score. Visualize confusion matrices and sample predictions. Reproducibility Notes Keep the same image size, data split ratio, and model parameters for consistent results. Use a fixed random seed to ensure reproducibility. Document your environment (Python version, library versions, hardware).
Institutions
- Daffodil International University