MPBD-18: A Large-Scale Real-World Medicinal Plant Image Dataset from Bangladesh for Automated Plant Species Identification

Published: 8 July 2026| Version 1 | DOI: 10.17632/26mb4gpzdn.1
Contributors:
,
,
,
,
,

Description

MPBD-18 is a large-scale real-world medicinal plant image dataset collected from diverse environments across Dhaka, Bangladesh, for automated species identification and computer vision research. The dataset contains 17,383 high-quality images belonging to 18 medicinal plant classes collected over a six-month period from November 2025 to May 2026. The dataset includes the following medicinal plant species: Aloe vera, Akondo, Amloki, Bashok, Bel, Bontulshi, Golmorich, Kalokeshi, Lojjaboti, Neem, Oporajita, Orjun, Osshogondha, Pathor Kuchi, Pudina, Shorpogondha, Telakucha, and Tulshi. MPBD-18 maintains a well-balanced class distribution, with approximately 900–1100 images per class to support reliable benchmark evaluation and machine learning applications. Images were collected from botanical gardens, medicinal plant gardens, plant nurseries, roadsides, open fields, and residential environments, including the National Botanical Garden, Dhaka International University Medicinal Plant Garden, and multiple nurseries located along 100 Feet Madani Avenue, Dhaka, Bangladesh. Images were captured using multiple smartphone devices under varying real-world conditions, including differences in illumination, background complexity, viewpoint, scale, and environmental settings. Burst-shot or continuous shutter mode images were intentionally avoided to maximize visual diversity and minimize redundant samples. The dataset was manually curated through duplicate removal, quality filtering, annotation verification, and preprocessing to ensure visual uniqueness and dataset reliability. The repository contains two primary components: (i) MPBD-18_Preprocessed_Dataset, which includes cleaned and normalized class-wise images, and (ii) MPBD-18_Benchmark_Split, which provides benchmark-ready training, validation, and testing subsets following an 80:10:10 split ratio with balanced class distributions to facilitate fair and reproducible model evaluation. Metadata and documentation files are also included. MPBD-18 is suitable for machine learning, deep learning, computer vision, biodiversity informatics, image classification, transfer learning, and automated medicinal plant recognition research. The dataset is intended to support reproducible benchmark evaluation and future research in automated medicinal plant recognition.

Files

Steps to reproduce

1. Download and extract the MPBD-18 dataset repository. 2. Access the MPBD-18_Preprocessed_Dataset directory for class-wise preprocessed images or use MPBD-18_Benchmark_Split for benchmark-ready training, validation, and testing subsets organized using an 80:10:10 split ratio with balanced class distributions. 3. Load the dataset using any standard machine learning or deep learning framework such as TensorFlow or PyTorch. 4. Use the metadata.csv file for image labels, scientific names, and split information. 5. Apply optional data augmentation during training if required for model development. 6. Train and evaluate image classification models using the provided benchmark split configuration.

Institutions

Categories

Artificial Intelligence, Computer Vision, Image Processing, Machine Learning, Medicinal Crop, Deep Learning, Agricultural Biotechnology

Licence