Preprocessed Image Dataset of Indian Tea Leaves for Plant Disease Recognition: Helopeltis and Red Spider based

Published: 14 September 2026| Version 2 | DOI: 10.17632/jm378tjnrt.2
Contributors:
,
,

Description

Present dataset contains a curated and preprocessed dataset of images of Indian tea leaves developed to facilitate research work in the field of plant disease recognition and automated classification using computer vision and machine learning techniques. The current dataset is dedicated to two important diseases of tea plants, i.e. Helopeltis infection and Red Spider and the healthy tea leaves, which constitute a three class classification problem under real field conditions. The dataset contains RGB images of the tea leaves from Indian regions of tea growing. Images were shot under natural light and background variations in order to maintain the complexity that occurs in the real world and between leaves in terms of leaf orientations, textures, and surface appearance. To improve the robustness and usability of deep learning models, extensive image preprocessing procedures were implemented and, as a result, several transformed versions of the original images were obtained. Three different classes are represented in the dataset: healthy tea leaves which have no known symptoms, tea leaves infested by Helopeltis with evident signs of puncture marks and localised necrotic areas, and tea leaves infested by Red Spider Mite with signs of bronzing, discoloration and surface degradation. Each class is stored as its own directory to enable the easy loading and labeling of models to be trained on the data. A total of ten preprocessing techniques were implemented on each of the original images to produce a variety of feature representations. These include conversing the image to grayscale, converting the image to binary, Gaussian blurring with kernel sizes of 3x3, 5x5, 7x7, and 9x9, adding Gaussian noise with intensities of 2% and 4%, and rotating the images with an angle of +15deg and -15deg. The inclusion of multiple preprocessing variants means that the influence of noise, texture, intensity and geometrical transformation on disease recognition performance can be studied. All of the images are standard formatted and arranged to ensure consistency between experiments. The dataset can be used for starting tasks such as image classification, impact analysis of preprocessing, feature extraction comparison and benchmarking of deep learning architectures. By providing both original visual features and a variety of preprocessed features, the dataset is favourable for the establishment of robust and generalizable models for plant disease detection. This dataset is published for academic and research purposes and is intended to contribute on precision agriculture, smart farming system/sustainable tea crop management based on data-driven disease diagnosis contribution.

Files

Steps to reproduce

The current dataset has been created through a structured image acquisition and preprocessing workflow; it has been designed to ensure transparency and reproducibility. Raw images of tea leaves were obtained from commercial tea plantations in Talap town, Assam, India, which is known to cultivate a lot of tea and experience repeated pest associated disease incidence. Image collection was conducted directly in the field in several gardens in the town to respect the natural variability of plant health, lighting conditions and morphology of the leaves. Approximately 1000 images of raw tea leaves were acquired by consumer grade digital cameras and smartphone devices. Images were captured under natural daylight with no artificial illumination to retain visual characteristics in the real world including shadows, background clutter and color variation. Leaves were photographed in different orientations, distances and angles, where no physical change or chemical treatment was made to the plants during the acquisition. Both leaves with no signs of damage and leaves with visible signs of Helopeltis infestation and Red Spider were recorded according to visual observation. Following the acquisition in the field, all raw images were transported to a controlled laboratory environment for preprocessing. Initial screening was done to eliminate blurred images, duplication or wrong exposures. The remaining images were organized in three categories, i.e., healthy, Helopeltis-affected, and Red Spider Mite-affected, according to observable disease symptoms. This manual labeling process was performed before the preprocessing process due to class ambiguity problem. Preprocessing of the image was done with standard image processing libraries in a Python-based environment. Each of raw images were downsized to have a uniform resolution for consistency when it came to training the model. A set of ten preprocessing operations were then applied independently for each image. These operations consisted of the grayscale conversion, binary thresholding, gaussian blurring with 3x3, 5x5, 7x7 and 9x9 kernel sizes, gaussian noise injection of 2% and 4% intensity levels, and rotational transformations of +15deg and 15deg in either direction. The preprocessing steps used deterministic parameters in all cases to enable the reproduction of results. The processed images were stored in lossless and high-quality formats and mirrored in class wise directory format, all with a consistent naming format that connects the processed image to the original image source.

Institutions

Categories

Computer Vision, Disease, Image Classification, Agricultural Plant, Tea

Licence