Medicinal Plant Leaf Disease Dataset

Published: 5 November 2025| Version 1 | DOI: 10.17632/ncg7kk3gwx.1
Contributors:
,
,
,

Description

This dataset contains 2,547 high-resolution leaf images (4000 × 3000 px, 4:3 aspect ratio) of three medicinal plant species, Kalanchoe pinnata (PatharKuchi), Azadirachta indica (Neem), and Ocimum tenuiflorum (Tulsi), collected from Senbag Upazila, Noakhali, Bangladesh, during August 2025. Each species includes one healthy and three disease or symptom classes, forming twelve balanced categories verified by agricultural experts. All images were captured in natural daylight using a 50 MP Honor 200 smartphone camera, ensuring clear color and texture representation. Images were taken from multiple angles and distances to preserve natural variations in illumination and background. Leaf samples were collected directly from local fields and home-grown medicinal gardens. Class distribution: Kalanchoe pinnata (PatharKuchi): Healthy (209), Web Blight (213), Yellow (204), Yellow Blight (213) → Total 839 Azadirachta indica (Neem): Healthy (244), Leaf Spot (211), Web Blight (205), Yellow (225) → Total 885 Ocimum tenuiflorum (Tulsi): Healthy (213), Downy Mildew (200), Web Blight (205), Yellow Spot (205) → Total 823 Grand Total: 2,547 images across 12 classes. Each image is stored in .jpg format (4000 × 3000 px) and named following a class-based convention (e.g., Neem_Healthy, Neem_LeafSpot, Tulsi_Healthy, Patharkuchi_YellowBlight). All files are placed in a single folder, with class names embedded in filenames for convenient parsing and automatic label extraction during model training. The dataset is designed for AI-based plant disease detection, particularly hybrid CNN/Transformer models and Explainable AI (XAI) research. It enables studies in agricultural image classification, medicinal plant health monitoring, and digital pathology applications. Verification and labeling were conducted under the supervision of the Upazila Agriculture Officer, Senbag, on 19 October 2025, ensuring correct disease identification and class validity.

Files

Steps to reproduce

Download and unzip the dataset. All images are stored in a single folder, each filename containing its class (e.g., Neem_LeafSpot.jpg). Use filename parsing to extract labels automatically. Resize images to 240×240 px or 224×224 px for model training. Split into train (70%), validation (15%), and test (15%) subsets. Apply light augmentation (rotation ±30°, brightness ±15%, horizontal flip). Train CNN, Transformer, or hybrid CNN–Transformer (LSeTNet) models. Evaluate using accuracy, precision, recall, and F1-score metrics. This dataset supports reproducible research in AI-driven medicinal plant disease classification, promoting open data practices in sustainable agriculture and plant pathology.

Institutions

  • Daffodil International University
  • Hamad Bin Khalifa University College of Science and Engineering

Categories

Agricultural Science, Computer Science, Artificial Intelligence, Computer Vision, Environmental Science, Plant Pathology, Medicinal and Aromatic Plants

Licence