Image Dataset on Chili Leaf Diseases in the Krishna River Basin of the Deccan Plateau, India
Description
Overview This dataset is a comprehensive resource for researchers and professionals in agriculture, machine learning, and computer vision, focusing on chili plant disease detection and growth stage classification. It provides high-resolution images of both healthy and diseased chili leaves, as well as samples representing various growth stages, making it highly suitable for deep learning and AI-based precision agriculture applications. Data Collection The images were systematically collected between October 2024 and February 2025 from chili plantations in the Bellary and Raichur districts of Karnataka, and the Prathipadu, Vatticherukuru, and Medikonduru of Guntur and Parchur, Yeddana Pudi, and MarturPrakasam districts of Andhra Pradesh, which are recognized as the global epicenters for red chili production. All field collection was conducted under the direct supervision of agricultural experts to ensure the pathological accuracy of the healthy and diseased leaf samples. The collection effort spanned across diverse agro-climatic zones of the Krishna River Basin within the Deccan Plateau. This multi-regional coverage ensured that a wide range of environmental and cultivation conditions were represented, capturing real-world variability in chili plant health, disease manifestation, and growth patterns. Using advanced digital cameras, high-resolution images were gathered, documenting different disease types and developmental stages to support robust AI-driven agricultural research. Dataset Structure The dataset is organized into two primary categories: Chili Leaf Disease Dataset This category includes images depicting six major leaf diseases affecting chili plants. Each image is meticulously labeled and verified by agricultural experts to enable precise classification and model training. Original Dataset: 1,856 high-resolution images (.jpg) Augmented Dataset: 12,000 enhanced images generated through data augmentation techniques (rotation, flipping, contrast adjustment, and zoom) # Disease Classes Covered: Bacterial Spot Curl Virus Cercospora Leaf Spot Nutrition Deficiency White Spot Healthy Leaves Specifically, this dataset enables: ✅ Early and Accurate Disease Detection — facilitating the identification of key foliar diseases such as Bacterial Spot, Curl Virus and Cercospora Leaf Spot at early stages, thereby preventing large-scale crop losses and improving disease management efficiency. ✅ Automated Growth Stage Monitoring — providing visual evidence of chili plants at various growth stages, enabling automated phenotyping, yield estimation and adaptive farm management strategies using AI and computer vision models. ✅ Benchmark for AI and Computer Vision Models — serving as a standardized dataset for training and benchmarking deep learning architectures such as CNNs, Vision Transformers and hybrid models for plant disease recognition, segmentation, and classification tasks.
Files
Steps to reproduce
Steps to Reproduce To reproduce the Chili Plant Leaf Disease Dataset from the Krishna River Basin, Deccan Plateau, India, follow the procedure outlined below. The process ensures scientific consistency, quality, and reproducibility for agricultural and computer vision research. 1. Site Selection: Select chili farms located across the Krishna River Basin, covering Bellary and Raichur (Karnataka) and Guntur, Prakasam, Krishna, and Kurnool (Andhra Pradesh) districts. These sites represent diverse climatic and soil conditions, allowing for natural variation in leaf disease expression. Permission from local farmers and agricultural officers should be obtained before data collection. 2. Image Acquisition: Images were collected between October 2024 and February 2025, using high-resolution digital cameras (≥12 MP) under natural daylight between 8:00–11:00 AM and 3:00–5:00 PM. Each image was captured approximately 20–30 cm from the leaf surface to ensure clarity. Both sides of leaves (adaxial and abaxial) were photographed, covering healthy, diseased, and nutrient-deficient samples across multiple growth stages (seedling to fruiting). 3. Annotation and Expert Validation: Agricultural experts manually classified the images into six categories: Bacterial Spot, Curl Virus, Cercospora Leaf Spot, Nutrition Deficiency, White Spot, and Healthy Leaves. Each label was verified through cross-validation among plant pathologists to ensure annotation accuracy and minimize labeling bias. 4. Image Preprocessing and Organization: Images were resized to uniform dimensions (256×256 or 512×512 pixels), normalized, and background noise was minimized using standard image processing tools. The dataset was systematically arranged into folders for training, validation, and testing. Metadata including capture date, location, and growth stage was recorded for each sample. 5. Data Augmentation: To enhance model generalization, augmentation techniques such as rotation, flipping, brightness adjustment, and zooming were applied, increasing the dataset from 1,856 original images to 12,000 enhanced samples. Augmentation parameters were optimized to preserve biological authenticity of disease patterns. 6. Dataset Validation and Benchmarking: The dataset was split into training (70%), validation (15%), and test (15%) subsets. Duplicate or corrupted images were removed. The dataset is suitable for use with deep learning frameworks such as TensorFlow, Keras, and PyTorch, and can be employed to train models like ResNet, DenseNet, or Vision Transformers for disease detection and classification. 7. Data Availability: For reproducibility, the dataset can be stored in open repositories such as IEEE DataPort, Zenodo, or Kaggle, accompanied by a README file, metadata, and annotation details to support future AI-driven agricultural research.
Institutions
- Manipal Academy of Higher EducationKarnataka, Manipal
- Manipal Institute of TechnologyKarnataka, Manipal
- Rao Bahadur Y Mahabaleshwarappa Engineering CollegeKarnataka, Bellary