Cauliflower Plant Health Image Dataset for Smart Agriculture Applications

Published: 14 January 2026| Version 1 | DOI: 10.17632/5f8zz9fr2p.1
Contributors:
Phani prasanthi parvathaneni, unnam Koteswara rao

Description

Dataset Description: This dataset comprises RGB images of cauliflower plants collected under real field conditions for the purpose of plant health assessment and smart agriculture research. The images were acquired from a local agricultural field located in Penamaluru village, near Vijayawada, Andhra Pradesh, India. Data collection was carried out during daytime under natural lighting conditions using a high-resolution mobile phone camera, ensuring realistic environmental variability. The dataset is organized into two distinct classes: Healthy and Damaged. The Healthy class contains images of visually healthy cauliflower plants without observable symptoms of disease or stress. The Damaged class includes images of plants exhibiting visible signs of pest infestation, disease, discoloration, or physiological stress. All images were manually reviewed and labeled to ensure accurate class representation.

Files

Steps to reproduce

Steps to Reproduce: 1. Download the ZIP file containing the cauliflower image dataset from Mendeley Data. 2. Extract the ZIP file to a local directory. 3. Open the main dataset folder, which contains two subfolders: Healthy and Damaged. 4. Images in the Healthy folder represent visually healthy cauliflower plants, while images in the Damaged folder represent diseased or stressed plants. 5. Load the images into an image-processing or machine learning environment such as MATLAB or Python (e.g., OpenCV, TensorFlow, or PyTorch). 6. Assign class labels based on the folder names. 7. Apply preprocessing techniques if required, such as resizing, normalization, or data augmentation. 8. Train and evaluate classification or detection models using standard machine learning workflows. Location of Data Collection: The images were collected from a local agricultural field located in Penamaluru village, near Vijayawada, Andhra Pradesh, India. The dataset represents real field conditions under natural lighting and environmental variability commonly observed in this agricultural region. Data Collection and Methodology: The dataset consists of RGB images of cauliflower plants collected from a local agricultural field under natural outdoor conditions. Image acquisition was carried out during daytime using a high-resolution mobile phone camera to ensure clear visualization of plant features. No artificial lighting or controlled laboratory environment was used, allowing the dataset to reflect real field conditions. Images were captured at multiple angles and distances to include variations in plant orientation, size, and visual appearance. Care was taken to include diverse samples representing different growth stages and visible plant health conditions. After acquisition, all images were manually reviewed and categorized into two classes based on visual inspection: Healthy and Damaged. The Healthy class contains images of cauliflower plants without visible disease or stress symptoms, while the Damaged class includes plants exhibiting signs of pest infestation, disease, discoloration, or physiological stress. The labeling was performed manually to ensure accurate class separation. The images were organized into a hierarchical folder structure, with separate directories for Healthy and Damaged classes. Standard JPG image format was used to maintain compatibility across platforms. Images were resized to a uniform resolution to ensure consistency for image processing and machine learning applications. No synthetic, augmented, or artificially generated images were included in the dataset.

Institutions

  • Pradad V Potluri Siddhartha Institute of Technology

Categories

Agricultural Plant

Licence