High-Resolution Dataset of Common Bean Leaf Diseases in Bangladesh

Published: 26 February 2026| Version 1 | DOI: 10.17632/89c9r5m2gx.1
Contributors:
,
,

Description

This dataset consists of 3,180 high-resolution images of Phaseolus vulgaris (common bean) leaves, collected from the Rangpur Division in Bangladesh. The dataset is divided into five classes: Angular Leaf Spot (ALS), Anthracnose (ANT), Mosaic Virus (MV), Rust, and Healthy Leaves. The images were captured under natural daylight conditions using mobile phones, ensuring real-world applicability for field-based disease detection. The dataset includes the following class distribution: Angular Leaf Spot (ALS): 635 images Anthracnose (ANT): 671 images Mosaic Virus (MV): 629 images Rust: 625 images Healthy Leaves: 620 images The dataset is well-balanced, which prevents bias during training, making it ideal for machine learning applications. The high-resolution images (average 4020 × 3093 pixels) offer detailed views of the leaf surface, essential for fine-grained feature extraction. The images were captured with an A4 background to minimize environmental disturbances and ensure consistent quality across the dataset. Accompanying the images is comprehensive metadata, which includes details on image resolution, file size, sharpness, noise, brightness, RGB values, and texture features like contrast and homogeneity. This metadata is useful for further analysis of disease-specific visual patterns and enables researchers to investigate disease classification under standardized conditions. This dataset is particularly suited for training and validating deep learning models such as Convolutional Neural Networks (CNNs), Vision Transformers, and YOLO-based architectures for disease classification and severity detection. It provides a valuable resource for AI-based plant pathology research, with potential applications in precision agriculture, aimed at reducing crop yield losses due to late disease detection. By enabling early disease identification, the dataset supports sustainable agricultural practices and enhances food security. The dataset is publicly available through Mendeley Data and comes in a ZIP archive containing the raw JPG images and corresponding CSV metadata files. The dataset is a reusable resource for machine learning-based plant disease detection, and it can be further augmented with other datasets for cross-regional benchmarking in disease monitoring and early intervention.

Files

Steps to reproduce

Field Data Collection: The dataset was collected from common bean (Phaseolus vulgaris) fields in the Rangpur Division of Bangladesh. The fields were selected based on their prevalence of common bean diseases and their relevance to the local farming community. Disease symptoms were observed on leaves of the beans, including Angular Leaf Spot, Anthracnose, Mosaic Virus, and Rust, as well as healthy leaves. These leaves were then carefully harvested from plants with local farmers' cooperation. Image Capture: Leaves were placed flat on a clean white A4 sheet (210 × 297 mm) to minimize environmental distractions such as soil and vegetation. Images were captured under ambient daylight conditions using consumer-grade mobile phones (e.g., Redmi-12 with a 50 MP camera). A smartphone camera was held approximately 30–40 cm from the leaf surface, perpendicular to the leaf to ensure clear, focused images. Multiple images were taken per leaf to ensure diversity in focus and angle. Only the highest-quality image from each leaf was retained. Disease Labeling: Each image was manually categorized into one of five classes based on the visual symptoms of the leaves: Angular Leaf Spot (ALS): Characterized by dark angular spots with necrotic centers. Anthracnose (ANT): Small, dark-bordered lesions found along leaf margins and veins. Mosaic Virus (MV): Mottled yellow-green and curled leaves. Rust: Orange-brown pustules on the underside of the leaves. Healthy Leaves: Leaves with no visible disease symptoms. Metadata Extraction: Automated metadata extraction was performed on all images using Python libraries such as OpenCV and scikit-image. The extracted metadata includes: Image resolution (in pixels), File size (in MB), Aspect ratio, Sharpness, brightness, and contrast measures, Noise level, RGB color values (mean per channel), and Texture features (contrast, homogeneity, energy, correlation using a gray-level co-occurrence matrix) This metadata provides detailed information about each image and is stored in CSV files, including comprehensive_statistics.csv, dataset Specifications.csv. Data Organization: The raw image files are organized into subfolders named after their respective classes. Each subfolder contains the corresponding images for that specific class. The dataset also includes CSV files for metadata and summary statistics. Visualization plots (e.g., bar charts, histograms) showing the distribution of file sizes, resolution, RGB values, and other quality features are included in PDF/PNG format. Quality Control: A thorough quality check was conducted to ensure that all images were high quality, with consistent focus, exposure, and minimal noise. The metadata for each image was cross-verified, ensuring that all relevant attributes were accurately recorded. The dataset was designed to minimize external noise from environmental factors and maintain high consistency across all images, enabling more effective use for machine learning model training.

Institutions

Categories

Artificial Intelligence, Computer Vision, Image Processing, Agricultural Engineering, Agricultural Health, Agriculture

Licence