Mahua Leaf Image Dataset

Published: 14 January 2026| Version 1 | DOI: 10.17632/4td9y2yg4v.1
Contributors:
Monica Sankat,

Description

The Mahua Leaf Image Dataset is a collection of high-resolution raw images of Mahua (Madhuca longifolia / Madhuca indica) plants captured under natural field conditions. The dataset has been developed under the NMICPS–TiHAN, IIT Hyderabad project, funded by the Department of Science and Technology (DST), Government of India, to support research in computer vision, federated learning, and drone-based precision agriculture. The current release comprises unannotated raw images collected in real agricultural environments, exhibiting varying illumination conditions, background clutter, soil textures, and natural leaf litter. The images capture Mahua plants at different growth stages and show visible variability in leaf color, texture, and stress symptoms, reflecting realistic field scenarios. This dataset is intended to serve as an initial open benchmark for research in plant health analysis, image classification, object detection, and decentralized (federated) learning frameworks. Expert-verified annotation and class-labeled versions (e.g., healthy, diseased, weed-affected) are planned as future updates. The current release comprises unannotated raw images collected in real agricultural environments, exhibiting varying illumination conditions, background clutter, soil textures, and natural leaf litter. The images capture Mahua plants at different growth stages and show visible variability in leaf color, texture, and stress symptoms, reflecting realistic field scenarios. This dataset is intended to serve as an initial open benchmark for research in plant health analysis, image classification, object detection, and decentralized (federated) learning frameworks. Expert-verified annotation and class-labeled versions (e.g., healthy, diseased, weed-affected) are planned as future updates.

Files

Steps to reproduce

Methodology: The dataset was created through on-site field visits to Mahua (Madhuca longifolia / Madhuca indica) growing locations, including agricultural fields and plantation areas. Data collection was performed under natural outdoor conditions without controlled environments. Instruments: Images were captured using commercially available DSLR cameras and smartphone cameras. No specialized sensors or laboratory equipment were used. Data Acquisition Protocol: Photographs were taken during daylight hours from multiple viewpoints, including full plant and leaf-level perspectives, allowing natural variations in illumination, shadows, soil background, and surrounding vegetation. Data Handling Workflow: Captured images were stored in standard image formats (JPEG) as original raw files, without cropping, enhancement, or annotation. Images were subsequently organized and sequentially renamed for dataset consistency. Reproducibility: The data collection process can be reproduced by conducting similar field visits and capturing Mahua plant images using standard consumer cameras under natural field conditions.

Institutions

  • VIT Bhopal University

Categories

Forestry, Computer Vision, Image Classification, Agroforestry, Medicinal Use of Plants, Plant Diseases

Funders

  • NMICPS TiHAN-IIT Hyderabad

Licence